github.com
OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.
Directory
Published directory entries matching your search.
OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.
Zero-configuration MCP server that unifies multiple AI coding assistants (Claude Code, Cursor, Codex) through intelligent auto-discovery.
Python-free Rust inference server with OpenAI API compatibility and hot model swapping.
Uniform deep learning inference framework for mobile, desktop and server.
Provides an optimized cloud and edge inferencing solution.
Agent skill and MCP server for tweet search, user lookup, follower export, media download, webhooks, and confirmation-gated X actions. MIT.