Benchmarks for what frontier AI hasn't solved
All Providers
Sep 03,2026
GPT-6 Astra
Leads Terminal-Bench 4.0 and Agents' Last Exam, strong at agentic coding.
RANK BY BENCHMARK
Agentic Coding 2.0
47.6%
#1/11
SWE-bench CLI
12.4%
#3/11
SnorkelFinance 2.0
11.6%
#7/11
SnorkelUnderwrite 2.0
21.2%
#5/10
SnorkelManufacturing
10%
#6/11
Sep 02,2026
Gemini Flash 3.8
Mid-pack across most boards, best showing on SnorkelLegal.
RANK BY BENCHMARK
Agentic Coding 2.0
32.6%
#5/11
SWE-bench CLI
6.3%
#4/11
SnorkelWorkplace
11.2%
#9/11
SnorkelFinance 2.0
5.4%
#10/11
SnorkelUnderwrite 2.0
19.1%
#8/10
Sep 02,2026
Muse Spark 1.3
Second on SnorkelLegal and Agents' Last Exam, weak on coding benchmarks.
RANK BY BENCHMARK
Agentic Coding 2.0
25%
#7/11
SWE-bench CLI
2.2%
#10/11
SnorkelWorkplace
17.7%
#1/11
SnorkelFinance 2.0
10%
#8/11
SnorkelUnderwrite 2.0
19.6%
#7/10
Sep 01,2026
Fable 5.1
Tops SWE-bench CLI, Terminal-Bench-Science, and Senior SWE-Bench.
RANK BY BENCHMARK
Agentic Coding 2.0
39.6%
#2/11
SWE-bench CLI
14.5%
#1/11
SnorkelWorkplace
15.8%
#3/11
SnorkelFinance 2.0
19.6%
#3/11
SnorkelUnderwrite 2.0
24%
#4/10
All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Software Engineering
Open Benchmarks Grants
Senior SWE-Bench
Evaluating coding agents on senior-level engineering work.
Tasteful Solve Rate
1
Fable 5.1
34.7%
2
Fable 5
34.7%
3
Opus 5
34.7%
Proprietary
Software Engineering
SWE-bench CLI
Tests whether an AI coding agent can diagnose and deliver a validated, multi-file change in a real open-source repository, navigating code, tests, dependencies, and tooling through the command line.
By pass@1
1
Fable 5.1
14.5%
2
Opus 5
14.0%
3
GPT-6 Astra
12.4%
Software Engineering
Open Benchmarks Grants
SlopCode Bench
Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.
Top Models by Iso Solve
1
GPT-5.5
28.06%
2
GPT-5.3-Codex
26.02%
3
GPT-5.4
23.47%
View archived benchmarks



