Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Capability/Efficiency
Open Benchmarks Grants
Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Sonnet 4.6  ·  ICL
Sonnet 4.6 · ICL
+0.196
2
GPT-5.4 ·  ICL
GPT-5.4 · ICL
+0.189
3
Sonnet 4.6  ·  Claude Code
Sonnet 4.6 · Claude Code
+0.185
Learn more about Continual Learning Bench
View archived benchmarks

For models that need to be right. Not just good enough.