SnorkelWorkplace
A frontier benchmark for evaluating AI agents on economically valuable professional work and verifiable workplace deliverables.
overview
WorkplaceAgents evaluates whether an agent can complete real-world professional tasks that require evidence synthesis, domain judgment, and a finished work product. Each task provides expert-authored instructions and relevant source materials, with success determined by the correctness and usefulness of the resulting deliverable.
Across its occupational coverage, the benchmark spans analytical, operational, technical, financial, scientific, administrative, healthcare, legal, and creative workflows. Tasks require multi-step reasoning and can produce documents, spreadsheets, presentations, code, and structured analyses.
At a glance
200
frontier tasks
19
sectors
96
occupations
19
resource formats
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Muse Spark 1.3 |
17.7%
|
30.3%
|
$10.4 | $51.99 |
| 2 | Grok 4.6 |
17.1%
|
28.6%
|
$7.84 | $39.19 |
| 3 | Fable 5.1 |
15.8%
|
30.8%
|
$18.65 | $93.27 |
| 4 | Qwen 3.8 Max |
14.9%
|
30.3%
|
$26.54 | $132.7 |
| 5 | Claude Opus 5 |
14.9%
|
34.7%
|
$15.96 | $79.79 |
| 6 | GPT-6 Astra |
14.4%
|
23.7%
|
$22.6 | $113.01 |
| 7 | Kimi K3 |
13%
|
28.2%
|
$7.79 | $38.96 |
| 8 | GLM 5.3 |
12.8%
|
23.6%
|
$14.33 | $71.64 |
| 9 | Gemini Flash 3.8 |
11.2%
|
23.7%
|
$3.63 | $18.15 |
| 10 | DeepSeek V4 Pro |
8.6%
|
35.2%
|
$1.61 | $8.04 |
| 11 | Nemotron 3 Ultra 550B |
6.6%
|
18.4%
|
$0.59 | $2.95 |
Want to evaluate your model against WorkplaceAgents? Talk to our team
Frontier performance
Methodology
Evaluator
Instructions, reference files, supporting materials, and required output formats are packaged with each task. Agents work through the tools available in the task environment.
scoring note
Scoring prioritizes substantive correctness, reasoning, and work-product quality. Formatting criteria are constrained so they cannot dominate the evaluation.
Behind the benchmark
Professional work rarely ends with a short answer. It requires finding relevant evidence, applying domain judgment, resolving ambiguity, and producing an artifact that another person can use.
WorkplaceAgents evaluates that end-to-end process. The benchmark measures whether an agent can produce a correct, useful, and professionally defensible work product across diverse forms of knowledge work.

