Chatbot Arena +

This leaderboard is based on the following benchmarks. Arena + — an agent-driven battle platform for large language models (LLMs). We use LLM-as-a-judge to compute Elo ratings. AAII — Artificial Analysis Intelligence Index aggregating 10 challenging evaluations. ARC-AGI — Artificial General Intelligence benchmark v2 to measure fluid intelligence.

Agent Harnesses

Curated, ranked list of AI agent harnesses — the runtimes that close the loop between a stateless model and the outside world. 150+ harnesses · 12 categories · weekly-rescored · MCP-ready : claude mcp add agent-harnesses -- uvx agent-harnesses-mcp

SWE-bench +

SWE-bench is a benchmark for evaluating large language models on real world software issues collected from GitHub. Given a codebase and an issue, a language model is tasked with generating a patch that resolves the described problem. SWE-bench Verified is a human-validated subset that more reliably evaluates AI models’ ability to solve issues. International Olympiad in Informatics (IOI) competition features standardized and automated grading.