Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Knowledge Work
Proprietary
WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Muse Spark 1.3
Muse Spark 1.3
17.7%
2
Grok 4.6
Grok 4.6
17.1%
3
Fable 5.1
Fable 5.1
15.8%
Learn more about WorkplaceAgents
Knowledge Work
Open Benchmarks Grants
Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
GPT-6 Astra
GPT-6 Astra
34.2
2
Muse Spark 1.3
Muse Spark 1.3
32.2%
3
Opus 5
Opus 5
31.6%
Learn more about Agents’ Last Exam
View archived benchmarks

For models that need to be right. Not just good enough.