Packrift optimization benchmark corpus generated from the top-1000 exact-spec feed, with public quality ledgers and operational benchmark pages.
Directory
Search results
Published directory entries matching your search.
Benchmarking Hallucination Detection Methods in RAG Towards Data Science provides machine-learning models, research, training resources, or…
Benchmarks of machine learning inference for Go.
Kaggle Benchmarks catalogs AI services and resources for browsing by feature or use case.
LLMs Bullshit Benchmark provides machine-learning models, research, training resources, or evaluation tools.
Wolfram LLM Benchmarking Project provides machine-learning models, research, training resources, or evaluation tools.
More human than human: measuring ChatGPT political bias Public Choice provides machine-learning models, research, training resources, or evaluation…
(Non-)Human — an art installation exploring the semi-human, semi-object territory.
Curated list of top Hugging Face models for NLP, vision, and audio tasks with demos and benchmarks.
GPT Migrate team is working on adding for the agent.
Codeflash uses AI to automatically find the most optimized version of your Python code through benchmarking — while verifying it's correct.
"a collaborative benchmark intended to probe large language models and extrapolate their future capabilities".
Library for hyperparameter optimization and black box optimization benchmarks.
Edit /src/streamlit app.py to customize this app to your heart's desire.:heart:.
A Challenging, Contamination-Free LLM Benchmark.
An open source robotics benchmark for meta- and multi-task reinforcement learning.
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Benchmarking Large Language Models.
CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
TruLens instruments your AI agent with OpenTelemetry, scores every step with benchmarked LLM judges, and tells you which version to ship.
Guidelines for Human-AI Interaction - Microsoft Research supports AI-assisted research, knowledge retrieval, summarization, or source analysis.
Not Human Search provides machine-learning models, research, training resources, or evaluation tools.
OutfitAnyone - a Hugging Face Space by HumanAIGC creates or edits images with generative AI and text-based controls.