OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.
Directory
Search results
Published directory entries matching your search.
Examples showing how to use the OpenAI vision API to run inference on images, video files and webcam streams.
Python-free Rust inference server with OpenAI API compatibility and hot model swapping.
AI-powered social listening tool that tracks brand mentions across Google, Instagram, Facebook, TikTok, Reddit, and X (Twitter). Uses HasData SERP API + LLM inference to extract structured insights, generate charts, export CSV reports, and send Telegram notifications on a 24-hour automated cycle.
Inference engine for TensorRT on Nvidia GPUs.
Uniform deep learning inference framework for mobile, desktop and server.
Provides an optimized cloud and edge inferencing solution.
Store, search, organize and make machine-learned inferences over big data at serving time.
A Whisper CLI client compatible with the original OpenAI client, using CTranslate2 for faster inference. opensource.
Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. (Archived).