Benchmarks for what frontier AI hasn't solved

All benchmarks
All Benchmarks
Sort: Newest
Benchmark Type: All
Proprietary
Enterprise Environments
SnorkelManufacturing

Measures how well AI agents can turn fragmented plant-floor, engineering, and supplier evidence into safe, technically defensible decisions and actions.

By pass@1
1
Grok 4.6
Grok 4.6
16.4%
2
GLM 5.3
GLM 5.3
11.9%
3
Fable 5.1
Fable 5.1
11.3%
Learn more about SnorkelManufacturing
Proprietary
Enterprise Environments
SnorkelRevOps

Tests whether AI agents can reconcile revenue systems, enforce hard commercial controls, and carry a decision through to the correct business state.

By pass@1
1
Grok 4.6
Grok 4.6
15.8%
2
Fable 5.1
Fable 5.1
14.0%
3
Kimi K3
Kimi K3
13.1%
Learn more about SnorkelRevOps
Proprietary
Enterprise Environments
SnorkelFinance 2.0

Scores how agents gather evidence, perform financial analysis, follow compliance constraints, and complete required state updates in a simulated environment.

By pass@1
1
Grok 4.6
Grok 4.6
25.1%
2
GLM 5.3
GLM 5.3
21.2%
3
 Fable 5.1
Fable 5.1
19.6%
Learn more about SnorkelFinance 2.0
Proprietary
Enterprise Environments
SnorkelUnderwrite 2.0

Measures an agent’s ability to turn incomplete, distributed insurance evidence into an auditable underwriting decision while respecting authority and policy constraints.

By pass@1
1
GLM 5.3
GLM 5.3
30.4%
2
DeepSeek V4 Pro
DeepSeek V4 Pro
28.5%
3
Kimi K3
Kimi K3
27.9%
Learn more about SnorkelUnderwrite 2.0
Proprietary
Enterprise Environments
SnorkelLegal

Evaluates AI agents on whether they can advance a legal matter while preserving the chain from controlling evidence to procedure, authority, and action.

By pass@1
1
Grok 4.6
Grok 4.6
40.7%
2
GLM 5.3
GLM 5.3
37.8%
3
Muse Spark 1.3
Muse Spark 1.3
36.5%
Learn more about SnorkelLegal
View archived benchmarks

For models that need to be right. Not just good enough.