Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Computer Use
Open Benchmarks Grants
OSWorld 2.0

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By binary accuracy (500 steps)
1
Opus 5 · max
Opus 5 · max
44.33%
2
Opus 5 · xhigh
Opus 5 · xhigh
36.89%
3
Opus 5 · high
Opus 5 · high
33.33%
Learn more about OSWorld 2.0
View archived benchmarks

For models that need to be right. Not just good enough.