Digital Work Index
Leaderboard
| Rank | Model | Score |
|---|---|---|
| 1 | Fable 5.1 |
23.62%
|
| 2 | GPT-6 Astra |
23.03%
|
| 3 | Opus 5 |
21.75%
|
| 4 | Grok 4.6 |
20.52%
|
| 5 | Gemini Flash 3.8 |
16.92%
|
| 6 | Kimi K3 |
16.76%
|
| 7 | GLM 5.3 |
16.45%
|
| 8 | Muse Spark 1.3 |
16.29%
|
| 9 | Qwen 3.8 Max |
11.81%
|
| 10 | DeepSeek V4 Pro |
11.07%
|
| Rank | Model | Score |
|---|---|---|
| 1 | GPT-6 Astra |
28.76%
|
| 2 | Fable 5.1 |
26.16%
|
| 3 | Opus 5 |
25.57%
|
| 4 | Grok 4.6 |
18.67%
|
| 5 | Gemini Flash 3.8 |
18.52%
|
| 6 | Kimi K3 |
14.32%
|
| 7 | GLM 5.3 |
13.44%
|
| 8 | Muse Spark 1.3 |
12.8%
|
| 9 | Qwen 3.8 Max |
9.45%
|
| 10 | DeepSeek V4 Pro |
8.11%
|
| Rank | Model | Score |
|---|---|---|
| 1 | GLM 5.3 |
27.02%
|
| 2 | Grok 4.6 |
26.93%
|
| 3 | Kimi K3 |
25.06%
|
| 4 | Fable 5.1 |
24.73%
|
| 5 | Muse Spark 1.3 |
22.16%
|
| 6 | DeepSeek V4 Pro |
21.2%
|
| 7 | Opus 5 |
19.01%
|
| 8 | Qwen 3.8 Max |
18.03%
|
| 9 | GPT-6 Astra |
16.85%
|
| 10 | Gemini Flash 3.8 |
16.03%
|
Methodology
Scope
Each benchmark is adjusted by the share of its mapped occupations employed in the simulated industry, using BLS industry-by-occupation data. Coding is treated as portable across industries. Specialized benchmarks are limited to their modeled industries. Shared occupations are split according to task concentration to avoid double counting. Each benchmark receives a wage term and a task-volume term. Its final weight is the normalized geometric mean of the two, giving equal influence to economic coverage and evaluated task volume.
Scoring and precision
Each model's score is the weighted average of its benchmark pass rates. With approximately 200 tasks per benchmark, estimated sampling precision is ±3.5 points per benchmark, giving a standard error of ±1.7 points on the index. Gaps between models smaller than roughly 2.4 points should not be treated as demonstrated ranking differences.
Behind the benchmark
The Digital Work Index summarizes model performance across eight challenging benchmarks covering software engineering, legal work, finance, insurance, manufacturing, and revenue operations. It combines results from Agentic Coding 2.0, SWE-bench CLI, SnorkelFinance 2.0, SnorkelUnderwrite 2.0, SnorkelManufacturing, SnorkelRevOps, SnorkelLegal, and Harvey LAB.
Each benchmark is weighted using two signals. The first is the wage base of the occupations it represents. The second is how much real-life work the benchmark reaches — measured by the jobs its tasks test. The current release covers ten models and 29 occupations. Together, the benchmarks represent 32.7% of the estimated $5.1 trillion annual wage bill across the 237 U.S. occupations classified here as digital work. The index provides an economically weighted comparison of model capability at the frontier. Productivity, automation, and labor-market impact are outside its scope.

