Language Benchmarks
Language Benchmarks is a developer resource for coding, documentation, communities, security research, or software tools.
Published directory entries matching your search.
Language Benchmarks is a developer resource for coding, documentation, communities, security research, or software tools.
Benchmarks Game is a developer resource for coding, documentation, communities, security research, or software tools.
Cybenetics PSU Benchmarks is a curated directory or link collection for discovering useful websites, tools, and resources.
Human Benchmark is a miscellaneous web resource for discovery, directories, archives, web culture, or useful tools.
Open Benchmarking is a miscellaneous web resource for discovery, directories, archives, web culture, or useful tools.
Packrift optimization benchmark corpus generated from the top-1000 exact-spec feed, with public quality ledgers and operational benchmark pages.
UNIGINE Benchmarks is a system utility for Windows, hardware checks, cleanup, process control, or desktop customization.
Benchmarking Hallucination Detection Methods in RAG Towards Data Science provides machine-learning models, research, training resources, or…
Benchmarks of machine learning inference for Go.
Kaggle Benchmarks catalogs AI services and resources for browsing by feature or use case.
LLMs Bullshit Benchmark provides machine-learning models, research, training resources, or evaluation tools.
Wolfram LLM Benchmarking Project provides machine-learning models, research, training resources, or evaluation tools.
Curated list of top Hugging Face models for NLP, vision, and audio tasks with demos and benchmarks.
Search basic in biology and medicine data collections and supporting dataset from CZE.
GPT Migrate team is working on adding for the agent.
Codeflash uses AI to automatically find the most optimized version of your Python code through benchmarking — while verifying it's correct.
"a collaborative benchmark intended to probe large language models and extrapolate their future capabilities".
Library for hyperparameter optimization and black box optimization benchmarks.
Edit /src/streamlit app.py to customize this app to your heart's desire.:heart:.
An open source robotics benchmark for meta- and multi-task reinforcement learning.
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.