Knowledge Work

SnorkelWorkplace

A frontier benchmark for evaluating AI agents on economically valuable professional work and verifiable workplace deliverables.

overview

WorkplaceAgents evaluates whether an agent can complete real-world professional tasks that require evidence synthesis, domain judgment, and a finished work product. Each task provides expert-authored instructions and relevant source materials, with success determined by the correctness and usefulness of the resulting deliverable.

Across its occupational coverage, the benchmark spans analytical, operational, technical, financial, scientific, administrative, healthcare, legal, and creative workflows. Tasks require multi-step reasoning and can produce documents, spreadsheets, presentations, code, and structured analyses.

At a glance

200

frontier tasks

19

sectors

96

occupations

19

resource formats

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
17.1%
28.6%
$7.84 $39.19
2 Fable 5.1
15.8%
30.8%
$18.65 $93.27
3 Opus 5
14.9%
34.7%
$15.96 $79.79
4 GPT-6 Astra
14.4%
23.7%
$22.6 $113.01
5 Kimi K3
13%
28.2%
$7.79 $38.96
6 GLM 5.3
12.8%
23.6%
$14.33 $71.64
7 Gemini Flash 3.8
11.2%
23.7%
$3.63 $18.15
8 DeepSeek V4 Pro
8.6%
35.2%
$1.61 $8.04
9 Nemotron 3 Ultra 550B
6.6%
18.4%
$0.59 $2.95
icon bulb

Want to evaluate your model against WorkplaceAgents? Talk to our team

Frontier performance

      Loading chart data...

      Methodology

      Evaluator

      Harbor evaluates completed work products using deterministic checks and rubric-based judging. Each rubric criterion maps to a substantive, independently verifiable requirement.
      timeout
      A run fails if the agent times out before producing a complete answer or required artifact. Each task is evaluated within a bounded execution environment.
      integration

      Instructions, reference files, supporting materials, and required output formats are packaged with each task. Agents work through the tools available in the task environment.

      scoring note

      Scoring prioritizes substantive correctness, reasoning, and work-product quality. Formatting criteria are constrained so they cannot dominate the evaluation.

      Behind the benchmark

      Professional work rarely ends with a short answer. It requires finding relevant evidence, applying domain judgment, resolving ambiguity, and producing an artifact that another person can use.

      WorkplaceAgents evaluates that end-to-end process. The benchmark measures whether an agent can produce a correct, useful, and professionally defensible work product across diverse forms of knowledge work.

      More benchmarks

      Agentic Coding

      Agentic Coding 2.0

      Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

      By pass@1
      1
      Image
      GPT-6 Astra
      47.6%
      2
      Image
      Fable 5.1
      39.6%
      3
      Image
      Opus 5
      38.9%
      Software Engineering

      SWE-bench CLI

      Evaluates validated multi-file code changes in open-source repositories.

      By pass@1
      1
      Image
      Fable 5.1
      14.5%
      2
      Image
      Opus 5
      14.0%
      3
      Image
      GPT-6 Astra
      12.4%
      Enterprise Environments

      SnorkelUnderwrite 2.0

      Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

      By pass@1
      1
      Image
      GLM 5.3
      30.4%
      2
      Image
      DeepSeek V4 Pro
      28.5%
      3
      Image
      Kimi K3
      27.9%
      Enterprise Environments

      SnorkelManufacturing

      Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

      By pass@1
      1
      Image
      Grok 4.6
      16.4%
      2
      Image
      GLM 5.3
      11.9%
      3
      Image
      Fable 5.1
      11.3%
      Enterprise Environments

      SnorkelRevOps

      Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

      By pass@1
      1
      Image
      Grok 4.6
      15.8%
      2
      Image
      Fable 5.1
      14.0%
      3
      Image
      Kimi K3
      13.1%
      Enterprise Environments

      SnorkelFinance 2.0

      Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

      By pass@1
      1
      Image
      Grok 4.6
      25.1%
      2
      Image
      GLM 5.3
      21.2%
      3
      Image
      Fable 5.1
      19.6%
      Enterprise Environments

      SnorkelLegal

      Evaluates legal workflows requiring evidence, procedure, authority, and action.

      By pass@1
      1
      Image
      Grok 4.6
      40.7%
      2
      Image
      GLM 5.3
      37.8%
      3
      Image
      Muse Spark 1.3
      36.5%
      Agentic Coding

      Terminal-Bench 4.0

      Evaluates terminal agents on continuously updated software-engineering tasks.

      By Resolution Rate
      1
      Image
      GPT-6 Astra
      58.2%
      2
      Image
      Fable 5.1
      57.9%
      3
      Image
      Opus 5
      53.9%
      Scientific & Research Workflows

      Terminal-Bench-Science

      Evaluates agents on scientific workflows derived from researchers’ own work.

      By Resolution Rate
      1
      Image
      Fable 5.1
      40.0%
      2
      Image
      Opus 5
      30.0%
      3
      Image
      GPT-5.6 Sol
      22.4%
      Agentic Coding

      Terminal-Bench 3.0

      Evaluates terminal agents on containerized tasks across seven domains.

      By Resolution Rate
      1
      Image
      Opus 5
      42.7%
      2
      Image
      GPT-5.6 Sol
      34.6%
      3
      Image
      Fable 5
      34.1%
      Computer Use

      OSWorld 2.0

      Evaluates computer-use agents on 108 long-horizon workflows.

      By binary accuracy (500 steps)
      1
      Image
      Opus 5 · max
      44.33%
      2
      Image
      Opus 5 · xhigh
      36.89%
      3
      Image
      Opus 5 · high
      33.33%
      Software Engineering

      Senior SWE-Bench

      Evaluates coding agents on senior-level software engineering tasks.

      Tasteful Solve Rate
      1
      Image
      Fable 5.1
      34.7%
      2
      Image
      Fable 5
      34.7%
      3
      Image
      Opus 5
      34.7%
      Knowledge Work

      Agents’ Last Exam

      Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

      By Binary Accuracy
      1
      Image
      GPT-6 Astra
      34.2%
      2
      Image
      Muse Spark 1.3
      32.2%
      3
      Image
      Opus 5
      31.6%
      Software Engineering

      SlopCode Bench

      Measures coding-agent performance across evolving software requirements.

      Top Models by Iso Solve
      1
      Image
      GPT-5.5
      28.06%
      2
      Image
      GPT-5.3-Codex
      26.02%
      3
      Image
      GPT-5.4
      23.47%
      Capability/Efficiency

      Continual Learning Bench

      Measures improvement across sequential, stateful tasks.

      Top Systems (Agg. Reward)
      1
      Image
      Sonnet 4.6 · ICL
      +0.196
      2
      Image
      GPT-5.4 · ICL
      +0.189
      3
      Image
      Sonnet 4.6 · Claude Code
      +0.185
      of

      For models that need to be right. Not just good enough.