Packrift optimization benchmark corpus generated from the top-1000 exact-spec feed, with public quality ledgers and operational benchmark pages.
Directory
Search results
Published directory entries matching your search.
Benchmarking Hallucination Detection Methods in RAG Towards Data Science provides machine-learning models, research, training resources, or…
EQ-Bench provides machine-learning models, research, training resources, or evaluation tools.
Benchmarks of machine learning inference for Go.
Kaggle Benchmarks catalogs AI services and resources for browsing by feature or use case.
LLMs Bullshit Benchmark provides machine-learning models, research, training resources, or evaluation tools.
Reward Bench Leaderboard - a Hugging Face Space by allenai provides machine-learning models, research, training resources, or evaluation tools.
Simple Bench provides machine-learning models, research, training resources, or evaluation tools.
Coding agents help developers plan, implement, review, test, and debug software. For independent capability comparisons, see SWE-bench and.
Wolfram LLM Benchmarking Project provides machine-learning models, research, training resources, or evaluation tools.
Simple yet flexible JavaScript charting library for the modern web.
Another Python Machine Learning Library.
Never miss another call with our AI-powered call answering service. 10x better than voicemail. 10x cheaper than a traditional phone answering service.
Generative AI will enrich investors and be deployed against everyone else.
Curated list of top Hugging Face models for NLP, vision, and audio tasks with demos and benchmarks.
GPT Migrate team is working on adding for the agent.
Codeflash uses AI to automatically find the most optimized version of your Python code through benchmarking — while verifying it's correct.
"a collaborative benchmark intended to probe large language models and extrapolate their future capabilities".
Library for hyperparameter optimization and black box optimization benchmarks.
Edit /src/streamlit app.py to customize this app to your heart's desire.:heart:.
A Challenging, Contamination-Free LLM Benchmark.
An open source robotics benchmark for meta- and multi-task reinforcement learning.
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Benchmarking Large Language Models.