An open source robotics benchmark for meta- and multi-task reinforcement learning.
Directory
Search results
Published directory entries matching your search.
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Website documenting commemorative benches.
Search an online collection of books, documents, media and archival records.
Benchmarking Large Language Models.
CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
Open-source project that supports software development with code generation, analysis, debugging, or documentation.
Here, CLI tools, libraries, Add-ons, Reports, Benchmarks and Sample Scripts for taking advantage of Google Apps Script which are publishing in my blog, Gists and GitHub are summarized.
TruLens instruments your AI agent with OpenTelemetry, scores every step with benchmarked LLM judges, and tells you which version to ship.
Research archive containing dark-web authorship verification datasets and baseline models.