Digital Work Index

The Digital Work Index is an economically weighted measure of model performance across eight, frontier-difficulty benchmarks in software engineering, logical reasoning and writing, and professional environments.

Leaderboard

Rank Model Score
1 Fable 5.1
23.62%
2 GPT-6 Astra
23.03%
3 Opus 5
21.75%
4 Grok 4.6
20.52%
5 Gemini Flash 3.8
16.92%
6 Kimi K3
16.76%
7 GLM 5.3
16.45%
8 Muse Spark 1.3
16.29%
9 Qwen 3.8 Max
11.81%
10 DeepSeek V4 Pro
11.07%
Rank Model Score
1 GPT-6 Astra
28.76%
2 Fable 5.1
26.16%
3 Opus 5
25.57%
4 Grok 4.6
18.67%
5 Gemini Flash 3.8
18.52%
6 Kimi K3
14.32%
7 GLM 5.3
13.44%
8 Muse Spark 1.3
12.8%
9 Qwen 3.8 Max
9.45%
10 DeepSeek V4 Pro
8.11%
Rank Model Score
1 GLM 5.3
27.02%
2 Grok 4.6
26.93%
3 Kimi K3
25.06%
4 Fable 5.1
24.73%
5 Muse Spark 1.3
22.16%
6 DeepSeek V4 Pro
21.2%
7 Opus 5
19.01%
8 Qwen 3.8 Max
18.03%
9 GPT-6 Astra
16.85%
10 Gemini Flash 3.8
16.03%

Methodology

Scope

We began with 825 BLS occupations and retain 237 whose work primarily produces digital outputs such as documents, datasets, software, designs, or decision records. Employment multiplied by BLS May 2025 mean annual wage defines the $5.105 trillion denominator out of $10.8 trillion tracked by BLS. The eight benchmarks map to 29 occupations.
Task mapping
Every evaluated task is assigned to an occupation. Seven benchmark mappings come from a comparison of task-level research to O*NET occupational taxonomies. Harvey LAB is mapped from its public task distribution.
Adjustments and weights

Each benchmark is adjusted by the share of its mapped occupations employed in the simulated industry, using BLS industry-by-occupation data. Coding is treated as portable across industries. Specialized benchmarks are limited to their modeled industries. Shared occupations are split according to task concentration to avoid double counting. Each benchmark receives a wage term and a task-volume term. Its final weight is the normalized geometric mean of the two, giving equal influence to economic coverage and evaluated task volume.

Scoring and precision

Each model's score is the weighted average of its benchmark pass rates. With approximately 200 tasks per benchmark, estimated sampling precision is ±3.5 points per benchmark, giving a standard error of ±1.7 points on the index. Gaps between models smaller than roughly 2.4 points should not be treated as demonstrated ranking differences.

Behind the benchmark

The Digital Work Index summarizes model performance across eight challenging benchmarks covering software engineering, legal work, finance, insurance, manufacturing, and revenue operations. It combines results from Agentic Coding 2.0, SWE-bench CLI, SnorkelFinance 2.0, SnorkelUnderwrite 2.0, SnorkelManufacturing, SnorkelRevOps, SnorkelLegal, and Harvey LAB.

Each benchmark is weighted using two signals. The first is the wage base of the occupations it represents. The second is how much real-life work the benchmark reaches — measured by the jobs its tasks test. The current release covers ten models and 29 occupations. Together, the benchmarks represent 32.7% of the estimated $5.1 trillion annual wage bill across the 237 U.S. occupations classified here as digital work. The index provides an economically weighted comparison of model capability at the frontier. Productivity, automation, and labor-market impact are outside its scope.

More benchmarks

Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Fable 5.1
39.6%
3
Image
Opus 5
38.9%
Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Fable 5.1
14.5%
2
Image
Opus 5
14.0%
3
Image
GPT-6 Astra
12.4%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
GLM 5.3
30.4%
2
Image
DeepSeek V4 Pro
28.5%
3
Image
Kimi K3
27.9%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
GLM 5.3
11.9%
3
Image
Fable 5.1
11.3%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Fable 5.1
19.6%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Opus 5
14.9%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Scientific & Research Workflows

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
Fable 5.1
40.0%
2
Image
Opus 5
30.0%
3
Image
GPT-5.6 Sol
22.4%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By binary accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · xhigh
36.89%
3
Image
Opus 5 · high
33.33%
Software Engineering

Senior SWE-Bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
Image
GPT-6 Astra
34.2%
2
Image
Muse Spark 1.3
32.2%
3
Image
Opus 5
31.6%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

Top Systems (Agg. Reward)
1
Image
Sonnet 4.6 · ICL
+0.196
2
Image
GPT-5.4 · ICL
+0.189
3
Image
Sonnet 4.6 · Claude Code
+0.185
of

For models that need to be right. Not just good enough.