LLMs Bullshit Benchmark provides machine-learning models, research, training resources, or evaluation tools.
Directory
Search results
Published directory entries matching your search.
Matin Libre publishes newspaper reporting, headlines and current-affairs coverage in Benin.
Nota Bene publishes newspaper reporting, headlines and current-affairs coverage in Slovakia.
O Benfica publishes newspaper reporting, headlines and current-affairs coverage in Portugal.
Rashtreeya Vidyalaya College of Engineering over in Bengaluru, India is pretty great for Energy, Physics and Astronomy and Computer Science.
Reward Bench Leaderboard - a Hugging Face Space by allenai provides machine-learning models, research, training resources, or evaluation tools.
Simple Bench provides machine-learning models, research, training resources, or evaluation tools.
Coding agents help developers plan, implement, review, test, and debug software. For independent capability comparisons, see SWE-bench and.
The fall of Pekin — scenes beneath the blood-stained walls of the Purple City. publishes newspaper reporting, headlines and current-affairs coverage.
Universita degli Studi del Sannio — Benevento, Italy — is strong in Pharmacology, Toxicology and Pharmaceutics.
Universite Sidi Mohamed Ben Abdellah in Fes, Morocco is a great pick for Mathematics, Business, Management and Accounting and Computer Science.
Wolfram LLM Benchmarking Project provides machine-learning models, research, training resources, or evaluation tools.
Curated list of top Hugging Face models for NLP, vision, and audio tasks with demos and benchmarks.
Search basic in biology and medicine data collections and supporting dataset from CZE.
GPT Migrate team is working on adding for the agent.
Fast Neural Networks framework built on top of Metal. Supports TensorFlow models.
Open-source platform for high-performance ML model serving.
Find published open datasets through filters, categories and keyword search.
Codeflash uses AI to automatically find the most optimized version of your Python code through benchmarking — while verifying it's correct.
Italian online music catalog/archive operated by Istituto centrale per i beni sonori ed audiovisivi.
"a collaborative benchmark intended to probe large language models and extrapolate their future capabilities".
Library for hyperparameter optimization and black box optimization benchmarks.
Edit /src/streamlit app.py to customize this app to your heart's desire.:heart:.
A Challenging, Contamination-Free LLM Benchmark.