labs.scale.com
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
Directory
Published directory entries matching your search.
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
Here, CLI tools, libraries, Add-ons, Reports, Benchmarks and Sample Scripts for taking advantage of Google Apps Script which are publishing in my blog, Gists and GitHub are summarized.
TruLens instruments your AI agent with OpenTelemetry, scores every step with benchmarked LLM judges, and tells you which version to ship.
Research archive containing dark-web authorship verification datasets and baseline models.