SnorkelManufacturing
Measures how well AI agents can turn fragmented plant-floor, engineering, and supplier evidence into safe, technically defensible decisions and actions.
overview
SnorkelManufacturing evaluates agents in a simulated discrete-parts plant, where engineering, production, quality, supplier, and safety decisions depend on evidence spread across operational records, reference documents, and specialized tools.
Tasks cover representative work such as CAD and tooling review, process and quality investigation, production planning, automation and safety checks, and work-instruction authoring. Agents must identify the relevant asset or revision, retrieve the evidence that actually applies, reconcile conflicting signals, and deliver a decision, artifact, or governed update.
The release concentrates on plant-floor decisions where engineering evidence, production pressure, and safety constraints collide.
At a glance
200
frontier tasks
33%
CAD tasks
37%
verifiable-output tasks
40.5%
tasks requiring state changes
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
16.4%
|
37.4%
|
$1.98 | $9.91 |
| 2 | GLM 5.3 |
11.9%
|
34.8%
|
$0.92 | $4.61 |
| 3 | Fable 5.1 |
11.3%
|
29.5%
|
$5.89 | $29.45 |
| 4 | Muse Spark 1.3 |
10.5%
|
33.5%
|
$1.09 | $5.46 |
| 5 | Kimi K3 |
10.2%
|
31.9%
|
$2.62 | $13.08 |
| 6 | GPT-6 Astra |
10%
|
23.7%
|
$7.39 | $36.94 |
| 7 | Claude Opus 5 |
9.4%
|
23.6%
|
$3.12 | $15.6 |
| 8 | Qwen 3.8 Max |
8.3%
|
24.4%
|
$2.23 | $11.16 |
| 9 | DeepSeek V4 Pro |
8.1%
|
25.4%
|
$1.08 | $5.42 |
| 10 | Gemini Flash 3.8 |
7.1%
|
19.9%
|
$1.05 | $5.25 |
| 11 | Nemotron 3 Ultra 550B |
5.4%
|
18.2%
|
$0.5 | $2.51 |
Want to evaluate your model against SnorkelUnderwrite 2.0? Talk to our team
Frontier performance
Methodology
Evaluator
The benchmark runs in a self-contained Docker environment using the OpenEnv interaction standard. Structured JSON MCP tools expose plant records, documents, engineering artifacts, analysis functions, and governed actions. Sessions preserve state, support seeded failure injection, and capture the full trajectory of user interaction, tool calls, and observations.
scoring note
A plausible explanation is not enough. An agent can fail by grounding its decision in the wrong asset or revision, ignoring a hard safety constraint, or leaving a required update incomplete. Valid alternative tool paths are allowed. Process and efficiency diagnostics expose unnecessary calls, tool errors, and failure to anchor conclusions in evidence.
Behind the benchmark
Plant-floor work is not a lookup problem. A supplier report can conflict with measured telemetry. A revision mismatch can invalidate an otherwise reasonable action. An urgent production request can create pressure to bypass a safety or change-control rule.
SnorkelManufacturing makes these conflicts part of the evaluation. It tests whether an agent can hold uncertain causes as uncertain, identify the evidence that controls the decision, respect safety boundaries, and carry the correct conclusion into the operating environment.
The subset combines CAD reasoning, checkable outcomes, and state-changing work so labs can distinguish technical fluency from dependable manufacturing operation.

