Directory

Search results

Published directory entries matching your search.

Search results

83 listings
github.com

An open source robotics benchmark for meta- and multi-task reinforcement learning.

Models & Machine Learning 251
github.com

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

Models & Machine Learning 279
openbenches.org

Website documenting commemorative benches.

Academic & Research Search 345
benyehuda.org

Search an online collection of books, documents, media and archival records.

Academic & Research Search 315
github.com

Benchmarking Large Language Models.

Models & Machine Learning 315
github.com

CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.

Models & Machine Learning 261
labs.scale.com

Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.

Models & Machine Learning 199
github.com

Open-source project that supports software development with code generation, analysis, debugging, or documentation.

AI Coding & Development 266

Here, CLI tools, libraries, Add-ons, Reports, Benchmarks and Sample Scripts for taking advantage of Google Apps Script which are publishing in my blog, Gists and GitHub are summarized.

Google Communities & Tools 283
trulens.org

TruLens instruments your AI agent with OpenTelemetry, scores every step with benchmarked LLM judges, and tells you which version to ship.

Models & Machine Learning 217
github.com

Research archive containing dark-web authorship verification datasets and baseline models.

Dark Web Archives & Historical Records 161