SnorkelRevOps
Tests whether AI agents can reconcile revenue systems, enforce hard commercial controls, and carry a decision through to the correct business state.
overview
SnorkelRevOps measures agents inside a simulated B2B SaaS revenue function. Work crosses CRM, CPQ, billing, compensation, marketing, usage, finance, and governance records, with each system authoritative for different decisions.
The release covers representative requests in forecasting, pricing and deal desk, territory and compensation, renewals, and revenue reconciliation. Agents must anchor to the right account or opportunity, ask for missing identifiers, reconcile competing records, apply proprietary controls, and either produce a decision-quality deliverable or execute the required update.
This public set concentrates the cross-system and control-sensitive tasks most likely to separate fluent assistants from reliable operators.
At a glance
200
frontier tasks
17
simulated personas represented
54%
semantic-output tasks
36.5%
tasks requiring state changes
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
15.8%
|
35%
|
$1.94 | $9.71 |
| 2 | Fable 5.1 |
14%
|
25.1%
|
$4.03 | $20.17 |
| 3 | Kimi K3 |
13.1%
|
31.5%
|
$1.77 | $8.85 |
| 4 | Muse Spark 1.3 |
12.6%
|
25.5%
|
$1.07 | $5.37 |
| 5 | Claude Opus 5 |
12.2%
|
23.9%
|
$4.02 | $20.11 |
| 6 | GLM 5.3 |
11.8%
|
25.7%
|
$1.27 | $6.36 |
| 7 | GPT-6 Astra |
10.9%
|
21.1%
|
$5.92 | $29.59 |
| 8 | Nemotron 3 Ultra 550B |
10.1%
|
26.5%
|
$0.34 | $1.68 |
| 9 | DeepSeek V4 Pro |
9.5%
|
24.1%
|
$0.78 | $3.91 |
| 10 | Qwen 3.8 Max |
6.5%
|
17.7%
|
$1.87 | $9.37 |
| 11 | Gemini Flash 3.8 |
5.4%
|
14.4%
|
$0.93 | $4.66 |
Want to evaluate your model against SnorkelRevOps? Talk to our team
Frontier performance
Methodology
Evaluator
Models operate through a stateful OpenEnv session backed by a reproducible Docker environment. Structured MCP calls connect the relevant revenue systems and document corpus. Simulated users control disclosure, tone, authority, and escalation. Every user exchange, tool call, observation, and state change is logged.
scoring note
Task correctness is gated on both the final answer and the required environment state. A model can be fluent and still fail if it cites the wrong source of truth, approves a prohibited exception, or executes a plan that diverges from the intended change. Valid alternate tool paths are not penalized. Redundant calls, tool errors, unsafe behavior, and plan-versus-action mismatches remain visible in diagnostics.
Behind the benchmark
RevOps is a key area of business where a numerically plausible answer can still cause operational harm. A discount can violate floor price. A forecast can mix CRM and billing numbers. A territory or compensation change can be valid in one system and wrong in another.
The benchmark requires agents to resolve source-of-truth, authority, and approval constraints before acting. It also tests whether they can manage incomplete information and human pressure without substituting generic business knowledge for the controls encoded in the environment.
The high semantic-output share measures judgment-quality communication, while state-changing tasks test operational follow-through. Simulated personas add the disclosure differences, escalation behavior, and authority boundaries that make revenue workflows difficult in practice.

