Cloud-Native LLM Routing Engine. Improve LLM app resilience and speed.
Directory
Search results
Published directory entries matching your search.
Open Bilingual Pre-Trained Model (ICLR 2023).
Open Bilingual Pre-Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs.
Implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.
Creating semantic cache to store responses from LLM queries.
GPU cluster manager for running and managing LLMs.
Create customizable UI components around your models.
A curated list of Large Language Model.
Python materials for the online course on diffusion models by @huggingface.
Platform for deploying your Machine Learning to production.
A service for deployment Apache Spark MLLib machine learning models as realtime, batch or reactive web services.
Code for hyperparameter tuning/optimization of machine learning and deep learning algorithms.
Open-source project that supports deploying, serving, monitoring, or operating AI and machine-learning systems.
Open-source project that supports deploying, serving, monitoring, or operating AI and machine-learning systems.
Kubernetes-based system for hyperparameter tuning and neural architecture search.
A python package that integrates an LLM copilot inside the keras model development workflow.
Kubernetes custom resource definition for serving ML models on arbitrary frameworks.
Open-source project that supports deploying, serving, monitoring, or operating AI and machine-learning systems.
Open-source project that supports deploying, serving, monitoring, or operating AI and machine-learning systems.
Standardized Serverless ML Inference Platform on Kubernetes.
Open-source all-in-one platform for engineering AI products. Traces, Evals, Datasets, Labels.
FastAPI framework to build production-grade LLM applications.
Developer-friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps!
Serverless LLM apps on Production with Jina AI Cloud (Archived).