Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Scientific & Research Workflows
Open Benchmarks Grants
Terminal-Bench-Science

A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

By Resolution Rate
1
Fable 5.1
Fable 5.1
40.0%
2
Opus 5
Opus 5
30.0%
3
GPT-5.6 Sol
GPT-5.6 Sol
22.4%
Learn more about Terminal-Bench-Science
View archived benchmarks

For models that need to be right. Not just good enough.