LLM Evaluation Metrics: Everything You Need for LLM Evaluation - Confident AI provides machine-learning models, research, training resources, or…
Directory
Search results
Published directory entries matching your search.
Falcon LLM is a generative large language model (LLM) that helps advance applications and use cases to future-proof our world.
LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.
Download this free LLM Testing Guide.
Custom Api is an AI chat or model resource for writing, coding, research, prompts, and assistant workflows.
Run LLM backends, APIs, frontends, and services with one command.
An LLM-based autonomous agent controlling real-world applications via RESTful APIs.
Minimize LLM token complexity to save API costs and model computations.
Aider LLM Leaderboards is an AI chat or model resource for writing, coding, research, prompts, and assistant workflows.
Indico LLM Leaderboard provides machine-learning models, research, training resources, or evaluation tools.
Llm Stats is an AI chat or model resource for writing, coding, research, prompts, and assistant workflows.
Patterns for Building LLM-based Systems & Products provides machine-learning models, research, training resources, or evaluation tools.
State of LLM Apps 2023 · Streamlit provides machine-learning models, research, training resources, or evaluation tools.
The Hugging Face Open LLM Leaderboard provides machine-learning models, research, training resources, or evaluation tools.
The Guide to LLM Evaluation Deci provides machine-learning models, research, training resources, or evaluation tools.
LLM Ops platform with Analytics, Monitoring, Evaluations and an LLM Optimization Studio powered by DSPy.
Documentation for LLM Evaluation Clarifai Guide, covering setup, features, and practical usage.
LLM Optimizer applies AI to data analysis, extraction, visualization, or business intelligence.
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
TensorZero builds open-source tools for production-grade LLM applications: LLM gateway, observability, optimization, evaluations, and experimentation.
Any URL to clean markdown for LLMs. Free API, no signup required. Strips JS/CSS and outputs LLM-ready content.
A Challenging, Contamination-Free LLM Benchmark.
andreasjansson/stable-diffusion-wip – Run with an API on Replicate provides machine-learning models, research, training resources, or evaluation…
APIDNA provides machine-learning models, research, training resources, or evaluation tools.