Open Bilingual Pre-Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs.
Directory
Search results
Published directory entries matching your search.
Google tool for building, testing, and experimenting with Gemini prompts, models, and API workflows.
Google tool for building, testing, and experimenting with Gemini prompts, models, and API workflows.
gotoHuman provides machine-learning models, research, training resources, or evaluation tools.
Implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.
Creating semantic cache to store responses from LLM queries.
Real-time GPU cloud price comparison across 30+ providers.
GPU cluster manager for running and managing LLMs.
Create customizable UI components around your models.
GraphPipe supports deploying, serving, monitoring, or operating AI and machine-learning systems.
GrimAC supports game modding, engines, custom content, or fan-game development workflows.
Groq supports deploying, serving, monitoring, or operating AI and machine-learning systems.
GTA5-Mods supports game modding, engines, custom content, or fan-game development workflows.
Guild AI supports deploying, serving, monitoring, or operating AI and machine-learning systems.
A curated list of Large Language Model.
Helicone AI supports deploying, serving, monitoring, or operating AI and machine-learning systems.
Hermes Agent helps build, test, automate, or manage AI agents, prompts, models, and API workflows.
HF Learn supports machine learning models, deployment, inspection, datasets, or AI development workflows.
Build, deploy, and scale production ML systems with Hopsworks. The Feature Store and MLOps platform for real-time AI, trusted by teams.
Python materials for the online course on diffusion models by @huggingface.
LLM evals platform for enterprises, providing tools to develop, evaluate, and observe AI systems.
Platform for deploying your Machine Learning to production.
A service for deployment Apache Spark MLLib machine learning models as realtime, batch or reactive web services.
Code for hyperparameter tuning/optimization of machine learning and deep learning algorithms.