OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.
Directory
Search results
Published directory entries matching your search.
Python-free Rust inference server with OpenAI API compatibility and hot model swapping.
Open-source DevOps agent to help you secure, deploy, and maintain production-ready infrastructure.
OpenTelemetry-based observability and monitoring for LLM and agents workflows.
Open-source project that supports deploying, serving, monitoring, or operating AI and machine-learning systems.
The open source standard for data logging. Enables ML monitoring and observability.
Bayesian Inference Tools in Python.
Deploy a ML inference service on a budget in less than 10 lines of code.
An easy-to-use feature store. Optimized for time-series data.
Fast inference engine for Transformer models in C++.
Curated list of awesome vector search framework/engine, library, cloud service and research papers to vector similarity search.
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
10x faster, cheaper, and better vector database.
Eurybia monitors data and model drift over time and securizes model deployment with data validation.
Interactive reports to analyze ML models during validation or production monitoring.
A library for efficient similarity search and clustering of dense vectors.
FeatherCNN is a high performance inference engine for convolutional neural networks.
An enterprise-grade, high performance feature store.
Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.
Go binding for MXNet c predict api to do inference with a pre-trained model.
Create customizable UI components around your models.
Build multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud Native.
Kubernetes-based system for hyperparameter tuning and neural architecture search.
Kubernetes custom resource definition for serving ML models on arbitrary frameworks.