Benchmarks for what frontier AI hasn't solved
All Providers
Sep 03,2026
GPT-6 Astra
Leads Terminal-Bench 4.0 and Agents' Last Exam, strong at agentic coding.
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/11
SWE-bench CLI
12.4%
#3/11
SnorkelFinance 2.0
11.6%
#7/11
SnorkelUnderwrite 2.0
21.2%
#5/10
SnorkelManufacturing
10%
#6/11
Sep 02,2026
Gemini Flash 3.8
Mid-pack across most boards, best showing on SnorkelLegal.
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#5/11
SWE-bench CLI
6.3%
#4/11
SnorkelWorkplace
11.2%
#9/11
SnorkelFinance 2.0
5.4%
#10/11
SnorkelUnderwrite 2.0
19.1%
#8/10
Sep 02,2026
Muse Spark 1.3
Second on SnorkelLegal and Agents' Last Exam, weak on coding benchmarks.
RANK BY BENCHMARK
Agentic Coding 2.0
25%
#7/11
SWE-bench CLI
2.2%
#10/11
SnorkelWorkplace
17.7%
#1/11
SnorkelFinance 2.0
10%
#8/11
SnorkelUnderwrite 2.0
19.6%
#7/10
Sep 01,2026
Fable 5.1
Tops SWE-bench CLI, Terminal-Bench-Science, and Senior SWE-Bench.
RANK BY BENCHMARK
Agentic Coding 2.0
39.6%
#2/11
SWE-bench CLI
14.5%
#1/11
SnorkelWorkplace
15.8%
#3/11
SnorkelFinance 2.0
19.6%
#3/11
SnorkelUnderwrite 2.0
24%
#4/10
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Scientific & Research Workflows
Open Benchmarks Grants
Terminal-Bench-Science
A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.
By Resolution Rate
1
Fable 5.1
40.0%
2
Opus 5
30.0%
3
GPT-5.6 Sol
22.4%
View archived benchmarks



