Falcon LLM is a generative large language model (LLM) that helps advance applications and use cases to future-proof our world.
Directory
Search results
Published directory entries matching your search.
Feast is an end-to-end open source feature store for machine learning. It allows teams to define, manage, discover, and serve features.
Feature store for machine learning.
Open-source project that provides machine-learning models, research, training resources, or evaluation tools.
FeatherCNN is a high performance inference engine for convolutional neural networks.
An enterprise-grade, high performance feature store.
Feature Stores for ML provides machine-learning models, research, training resources, or evaluation tools.
A Virtual Feature Store. Turn your existing data infrastructure into a feature store.
Open-source project that provides machine-learning models, research, training resources, or evaluation tools.
PyTorch Lightning extension that accelerates and enhances foundation model experimentation with flexible fine-tuning schedules.
Ultravox is a real-time voice AI infrastructure layer that powers fast, natural, and scalable voice agents.
Ultravox is a real-time voice AI infrastructure layer that powers fast, natural, and scalable voice agents.
Running large language models on a single GPU for throughput-oriented scenarios. (Archived).
Drag & drop UI to build your customized LLM flow using LangchainJS.
Library for high performance deep learning inference on NVIDIA GPUs.
Open-source self-hostable end-to-end LLMOps platform unifying tracing, evals, simulations, datasets, gateway, and guardrails.
Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.
Testing framework dedicated to ML models, from tabular to LLMs. Detect risks of biases, performance issues and errors in 4 lines of code.
Cloud-Native LLM Routing Engine. Improve LLM app resilience and speed.
Open Bilingual Pre-Trained Model (ICLR 2023).
Open Bilingual Pre-Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs.
gotoHuman provides machine-learning models, research, training resources, or evaluation tools.
Implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.
Creating semantic cache to store responses from LLM queries.