Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Agentic Coding
Open Benchmarks Grants
Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Opus 5
Opus 5
42.7%
2
GPT-5.6 Sol
GPT-5.6 Sol
34.6%
3
Fable 5
Fable 5
34.1%
Learn more about Terminal-Bench 3.0
Proprietary
Agentic Coding
Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
GPT-6 Astra
GPT-6 Astra
47.6%
2
Fable 5.1
Fable 5.1
39.6%
3
Opus 5
Opus 5
38.9%
Learn more about Agentic Coding 2.0
Agentic Coding
Open Benchmarks Grants
Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
GPT-6 Astra
GPT-6 Astra
58.2%
2
Fable 5.1
Fable 5.1
57.9%
3
Opus 5
Opus 5
53.9%
Learn more about Terminal-Bench 4.0
View archived benchmarks

For models that need to be right. Not just good enough.