Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Software Engineering
Open Benchmarks Grants
Senior SWE-Bench

Evaluating coding agents on senior-level engineering work.

Tasteful Solve Rate
1
Fable 5.1
Fable 5.1
34.7%
2
Fable 5
Fable 5
34.7%
3
Opus 5
Opus 5
34.7%
Learn more about Senior SWE-Bench
Proprietary
Software Engineering
SWE-bench CLI

Tests whether an AI coding agent can diagnose and deliver a validated, multi-file change in a real open-source repository, navigating code, tests, dependencies, and tooling through the command line.

By pass@1
1
Fable 5.1
Fable 5.1
14.5%
2
Opus 5
Opus 5
14.0%
3
GPT-6 Astra
GPT-6 Astra
12.4%
Learn more about SWE-bench CLI
Software Engineering
Open Benchmarks Grants
SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
GPT-5.5
GPT-5.5
28.06%
2
GPT-5.3-Codex
GPT-5.3-Codex
26.02%
3
GPT-5.4
GPT-5.4
23.47%
Learn more about SlopCode Bench
View archived benchmarks

For models that need to be right. Not just good enough.