Platform for deploying your Machine Learning to production.
Directory
Search results
Published directory entries matching your search.
A service for deployment Apache Spark MLLib machine learning models as realtime, batch or reactive web services.
Code for hyperparameter tuning/optimization of machine learning and deep learning algorithms.
A python package that integrates an LLM copilot inside the keras model development workflow.
FastAPI framework to build production-grade LLM applications.
Serverless LLM apps on Production with Jina AI Cloud (Archived).
LLM Ops platform with Analytics, Monitoring, Evaluations and an LLM Optimization Studio powered by DSPy.
LLM App is a Python library that helps you build real-time AI-powered data pipelines with few lines of code.
A tool that provides a Model Context Protocol (MCP) server for any web application with an OpenAPI specification.
A chat LLM based on the state-space model architecture.
A model-agnostic visual debugging tool for machine learning.
Small Language Model tailored for edge devices.
A platform for deploying and serving machine learning models.
MindSpore is a new open source deep learning training/inference framework that could be used for mobile, edge and cloud scenarios.
Intuitive convenience tooling for lightning-fast, efficient development and ensuring quality in LLM-based applications.
OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others).
Serve Llama 2 and other large language models locally from command line or through a browser interface.
Fujitsu Research's post-training quantization pipeline for LLMs (QEP, AutoBit, JointQ, rotation) with vLLM plugin (arXiv:2603.28845).
Jan - Run LLMs like Mistral or Llama2 locally and offline on your computer, or connect to remote AI APIs.
Use AutoML to do model compression.
A high-speed inference engine for deploying LLMs locally.
Open-source tool to simplify the process of creating and managing LLM workflows and prompts as a self-hosted solution.
AI-generated visualization prototyping and editing platform, support 2D, 3D models, combined with LLM(Large Language Model) for quick editing.
Inference engine for TensorRT on Nvidia GPUs.