Open-source tool that lets you package ML models in a standard, production-ready container.
Directory
Search results
Published directory entries matching your search.
Fast inference engine for Transformer models in C++.
Curated list of awesome vector search framework/engine, library, cloud service and research papers to vector similarity search.
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
A library for probabilistic modelling, inference, and criticism. Built on top of TensorFlow.
10x faster, cheaper, and better vector database.
Eurybia monitors data and model drift over time and securizes model deployment with data validation.
A library for efficient similarity search and clustering of dense vectors.
FeatherCNN is a high performance inference engine for convolutional neural networks.
An enterprise-grade, high performance feature store.
A Virtual Feature Store. Turn your existing data infrastructure into a feature store.
Ultravox is a real-time voice AI infrastructure layer that powers fast, natural, and scalable voice agents.
Ultravox is a real-time voice AI infrastructure layer that powers fast, natural, and scalable voice agents.
Open-source self-hostable end-to-end LLMOps platform unifying tracing, evals, simulations, datasets, gateway, and guardrails.
Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.
Go binding for MXNet c predict api to do inference with a pre-trained model.
Real-time GPU cloud price comparison across 30+ providers.
Create customizable UI components around your models.
Build, deploy, and scale production ML systems with Hopsworks. The Feature Store and MLOps platform for real-time AI, trusted by teams.
Helps teams build, deploy, observe or operate machine-learning systems.
Build multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud Native.
Kubernetes-based system for hyperparameter tuning and neural architecture search.
Respan unifies LLM observability, evals, prompt optimization, and an AI gateway so teams can ship reliable AI applications.
Kubernetes custom resource definition for serving ML models on arbitrary frameworks.