Benchmarks for what frontier AI hasn't solved
All Providers
Sep 03,2026
GPT-6 Astra
Leads Terminal-Bench 4.0 and Agents' Last Exam, strong at agentic coding.
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/11
SWE-bench CLI
12.4%
#3/11
SnorkelFinance 2.0
11.6%
#7/11
SnorkelUnderwrite 2.0
21.2%
#5/10
SnorkelManufacturing
10%
#6/11
Sep 02,2026
Gemini Flash 3.8
Mid-pack across most boards, best showing on SnorkelLegal.
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#5/11
SWE-bench CLI
6.3%
#4/11
SnorkelWorkplace
11.2%
#9/11
SnorkelFinance 2.0
5.4%
#10/11
SnorkelUnderwrite 2.0
19.1%
#8/10
Sep 02,2026
Muse Spark 1.3
Second on SnorkelLegal and Agents' Last Exam, weak on coding benchmarks.
RANK BY BENCHMARK
Agentic Coding 2.0
25%
#7/11
SWE-bench CLI
2.2%
#10/11
SnorkelWorkplace
17.7%
#1/11
SnorkelFinance 2.0
10%
#8/11
SnorkelUnderwrite 2.0
19.6%
#7/10
Sep 01,2026
Fable 5.1
Tops SWE-bench CLI, Terminal-Bench-Science, and Senior SWE-Bench.
RANK BY BENCHMARK
Agentic Coding 2.0
39.6%
#2/11
SWE-bench CLI
14.5%
#1/11
SnorkelWorkplace
15.8%
#3/11
SnorkelFinance 2.0
19.6%
#3/11
SnorkelUnderwrite 2.0
24%
#4/10
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Capability/Efficiency
Open Benchmarks Grants
Continual Learning Bench
Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.
Top Systems (Agg. Reward)
1
Sonnet 4.6 · ICL
+0.196
2
GPT-5.4 · ICL
+0.189
3
Sonnet 4.6 · Claude Code
+0.185
View archived benchmarks



