Inference for text-embedding models.
Directory
Search results
Published directory entries matching your search.
Large Language Model Text Generation Inference.
A machine learning / bayesian inference assigning attributes to objects.
Drop-in AsyncOpenAI replacement that transparently batches requests via the Batch API for cheaper LLM inference.
Bayesian Inference Tools in Python.
Official inference framework for 1-bit LLMs, by Microsoft. opensource.
Deploy a ML inference service on a budget in less than 10 lines of code.
An agentic company research tool powered by LangGraph and Tavily that conducts deep diligence on companies using a multi-agent framework. It leverages Google's Gemini 2.5 Flash and OpenAI's GPT-5.1 on the backend for inference.
Fast inference engine for Transformer models in C++.
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
Fast inference engine for whisper in C++ using CTranslate2.
FeatherCNN is a high performance inference engine for convolutional neural networks.
Library for high performance deep learning inference on NVIDIA GPUs.
Benchmarks of machine learning inference for Go.
Go binding for MXNet c predict api to do inference with a pre-trained model.
Standardized Serverless ML Inference Platform on Kubernetes.
Neural network inference from the command line, implemented in CHICKEN Scheme.
A lightweight, portable pure C99 onnx inference engine for embedded devices with hardware acceleration support.
Extensible Toolkit for Finetuning and Inference of Large Foundation Models.
MindSpore is a new open source deep learning training/inference framework that could be used for mobile, edge and cloud scenarios.
Ncnn is a high-performance neural network inference framework optimized for the mobile platform.
Easy-to-use library to boost AI inference.
Neural networks framework in pure C: training and inference, no dependencies.
A high-speed inference engine for deploying LLMs locally.