LLM Evaluation Metrics: Everything You Need for LLM Evaluation - Confident AI provides machine-learning models, research, training resources, or…
Directory
Search results
Published directory entries matching your search.
Falcon LLM is a generative large language model (LLM) that helps advance applications and use cases to future-proof our world.
Aider LLM Leaderboards is an AI chat or model resource for writing, coding, research, prompts, and assistant workflows.
Indico LLM Leaderboard provides machine-learning models, research, training resources, or evaluation tools.
Documentation for LLM Evaluation Clarifai Guide, covering setup, features, and practical usage.
LLM Optimizer applies AI to data analysis, extraction, visualization, or business intelligence.
Llm Stats is an AI chat or model resource for writing, coding, research, prompts, and assistant workflows.
Download this free LLM Testing Guide.
LLM Visualization is an AI chat or model resource for writing, coding, research, prompts, and assistant workflows.
Patterns for Building LLM-based Systems & Products provides machine-learning models, research, training resources, or evaluation tools.
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
State of LLM Apps 2023 · Streamlit provides machine-learning models, research, training resources, or evaluation tools.
TensorZero builds open-source tools for production-grade LLM applications: LLM gateway, observability, optimization, evaluations, and experimentation.
The Hugging Face Open LLM Leaderboard provides machine-learning models, research, training resources, or evaluation tools.
The Guide to LLM Evaluation Deci provides machine-learning models, research, training resources, or evaluation tools.
Wiki LLM List is a Wikipedia reference page for background information, comparisons, and linked sources.
An introduction to the concepts behind AgentGPT, BabyAGI, LangChain, and the LLM-powered agent revolution.
Uncover LLM evaluation's importance and explore methods for assessing its performance and impact across industries.
Collection of patterns for experimenting with agents, llm pipelines, and ChainOfThoughtStrategy.
Open source generative AI development platform for building AI agents, LLM orchestration, and more.
An LLM by xAI with open source and open weights. opensource.
Indic LLM Arena provides machine-learning models, research, training resources, or evaluation tools.
Respan unifies LLM observability, evals, prompt optimization, and an AI gateway so teams can ship reliable AI applications.
A Challenging, Contamination-Free LLM Benchmark.