Enterprise Environments

SnorkelRevOps

Tests whether AI agents can reconcile revenue systems, enforce hard commercial controls, and carry a decision through to the correct business state.

overview

SnorkelRevOps measures agents inside a simulated B2B SaaS revenue function. Work crosses CRM, CPQ, billing, compensation, marketing, usage, finance, and governance records, with each system authoritative for different decisions.

The release covers representative requests in forecasting, pricing and deal desk, territory and compensation, renewals, and revenue reconciliation. Agents must anchor to the right account or opportunity, ask for missing identifiers, reconcile competing records, apply proprietary controls, and either produce a decision-quality deliverable or execute the required update.

This public set concentrates the cross-system and control-sensitive tasks most likely to separate fluent assistants from reliable operators.

At a glance

200

frontier tasks

17

simulated personas represented

54%

semantic-output tasks

36.5%

tasks requiring state changes

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
15.8%
35%
$1.94 $9.71
2 Fable 5.1
14%
25.1%
$4.03 $20.17
3 Kimi K3
13.1%
31.5%
$1.77 $8.85
4 Muse Spark 1.3
12.6%
25.5%
$1.07 $5.37
5 Claude Opus 5
12.2%
23.9%
$4.02 $20.11
6 GLM 5.3
11.8%
25.7%
$1.27 $6.36
7 GPT-6 Astra
10.9%
21.1%
$5.92 $29.59
8 Nemotron 3 Ultra 550B
10.1%
26.5%
$0.34 $1.68
9 DeepSeek V4 Pro
9.5%
24.1%
$0.78 $3.91
10 Qwen 3.8 Max
6.5%
17.7%
$1.87 $9.37
11 Gemini Flash 3.8
5.4%
14.4%
$0.93 $4.66
icon bulb

Want to evaluate your model against SnorkelRevOps? Talk to our team

Frontier performance

Loading chart data...

Methodology

Evaluator

Numeric and structured outputs are verified programmatically. Memos, forecasts, adjudications, and other semantic outputs are evaluated for meaning against scenario-specific rubrics. State hashes verify required updates. Additional judges assess policy-safe communication, tool use, user interaction, and efficiency.
timeout
A fixed run budget governs time and agent steps so multi-tool execution can be compared across models. Results should be reported with the timeout and step configuration used for the evaluation.
integration

Models operate through a stateful OpenEnv session backed by a reproducible Docker environment. Structured MCP calls connect the relevant revenue systems and document corpus. Simulated users control disclosure, tone, authority, and escalation. Every user exchange, tool call, observation, and state change is logged.

scoring note

Task correctness is gated on both the final answer and the required environment state. A model can be fluent and still fail if it cites the wrong source of truth, approves a prohibited exception, or executes a plan that diverges from the intended change. Valid alternate tool paths are not penalized. Redundant calls, tool errors, unsafe behavior, and plan-versus-action mismatches remain visible in diagnostics.

Behind the benchmark

RevOps is a key area of business where a numerically plausible answer can still cause operational harm. A discount can violate floor price. A forecast can mix CRM and billing numbers. A territory or compensation change can be valid in one system and wrong in another.

The benchmark requires agents to resolve source-of-truth, authority, and approval constraints before acting. It also tests whether they can manage incomplete information and human pressure without substituting generic business knowledge for the controls encoded in the environment.

The high semantic-output share measures judgment-quality communication, while state-changing tasks test operational follow-through. Simulated personas add the disclosure differences, escalation behavior, and authority boundaries that make revenue workflows difficult in practice.

More benchmarks

Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Fable 5.1
39.6%
3
Image
Opus 5
38.9%
Software Engineering

SWE-bench CLI

Tests whether an AI coding agent can diagnose and deliver a validated, multi-file change in a real open-source repository, navigating code, tests, dependencies, and tooling through the command line.

By pass@1
1
Image
Fable 5.1
14.5%
2
Image
Opus 5
14.0%
3
Image
GPT-6 Astra
12.4%
Enterprise Environments

SnorkelUnderwrite 2.0

Measures an agent’s ability to turn incomplete, distributed insurance evidence into an auditable underwriting decision while respecting authority and policy constraints.

By pass@1
1
Image
GLM 5.3
30.4%
2
Image
DeepSeek V4 Pro
28.5%
3
Image
Kimi K3
27.9%
Enterprise Environments

SnorkelManufacturing

Measures how well AI agents can turn fragmented plant-floor, engineering, and supplier evidence into safe, technically defensible decisions and actions.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
GLM 5.3
11.9%
3
Image
Fable 5.1
11.3%
Enterprise Environments

SnorkelFinance 2.0

Scores how agents gather evidence, perform financial analysis, follow compliance constraints, and complete required state updates in a simulated environment.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Fable 5.1
19.6%
Enterprise Environments

SnorkelLegal

Evaluates AI agents on whether they can advance a legal matter while preserving the chain from controlling evidence to procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Muse Spark 1.3
17.7%
2
Image
Grok 4.6
17.1%
3
Image
Fable 5.1
15.8%
Agentic Coding

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Scientific & Research Workflows

Terminal-Bench-Science

A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

By Resolution Rate
1
Image
Fable 5.1
40.0%
2
Image
Opus 5
30.0%
3
Image
GPT-5.6 Sol
22.4%
Agentic Coding

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Computer Use

OSWorld 2.0

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By binary accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · xhigh
36.89%
3
Image
Opus 5 · high
33.33%
Software Engineering

Senior SWE-Bench

Evaluating coding agents on senior-level engineering work.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
Image
GPT-6 Astra
34.2
2
Image
Muse Spark 1.3
32.2%
3
Image
Opus 5
31.6%
Software Engineering

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
Sonnet 4.6 · ICL
+0.196
2
Image
GPT-5.4 · ICL
+0.189
3
Image
Sonnet 4.6 · Claude Code
+0.185
of

For models that need to be right. Not just good enough.