Benchmarks for what frontier AI hasn't solved
Our ability to measure AI has been outpaced by our ability to develop it. We close that evaluation gap with coding benchmarks built around the tasks today's agents still break down on.
Senior SWE-Bench
A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.


Terminal-Bench 3.0
The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.
Agents’ Last Exam
Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.
SlopCode Bench
Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.
Terminal-Bench 2.1
Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.
Agentic Coding
A benchmark for evaluating AI models on complex, real-world coding tasks that require multi-step reasoning, tool use, and autonomous problem-solving.
Learn more
A curated collection of blogs, research papers, and reading-group discussions on agentic coding benchmarks, covering benchmark design, realistic coding environments, model performance, and the failure modes shaping what comes next.









