Language Benchmarks is a developer resource for coding, documentation, communities, security research, or software tools.
Directory
Search results
Published directory entries matching your search.
Benchmarks Game is a developer resource for coding, documentation, communities, security research, or software tools.
Cybenetics PSU Benchmarks is a curated directory or link collection for discovering useful websites, tools, and resources.
Human Benchmark is a miscellaneous web resource for discovery, directories, archives, web culture, or useful tools.
Open Benchmarking is a miscellaneous web resource for discovery, directories, archives, web culture, or useful tools.
Packrift optimization benchmark corpus generated from the top-1000 exact-spec feed, with public quality ledgers and operational benchmark pages.
UNIGINE Benchmarks is a system utility for Windows, hardware checks, cleanup, process control, or desktop customization.
Benchmarking Hallucination Detection Methods in RAG Towards Data Science provides machine-learning models, research, training resources, or…
Kaggle Benchmarks catalogs AI services and resources for browsing by feature or use case.
LLMs Bullshit Benchmark provides machine-learning models, research, training resources, or evaluation tools.
Wolfram LLM Benchmarking Project provides machine-learning models, research, training resources, or evaluation tools.
Search basic in biology and medicine data collections and supporting dataset from CZE.
Codeflash uses AI to automatically find the most optimized version of your Python code through benchmarking — while verifying it's correct.
Edit /src/streamlit app.py to customize this app to your heart's desire.:heart:.
A Challenging, Contamination-Free LLM Benchmark.
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
TruLens instruments your AI agent with OpenTelemetry, scores every step with benchmarked LLM judges, and tells you which version to ship.