SnorkelFinance 2.0
Scores how agents gather evidence, perform financial analysis, follow compliance constraints, and complete required state updates in a simulated environment.
overview
SnorkelFinance 2.0 evolves the original SnorkelFinance benchmark from report-grounded financial QA into an environment-based evaluation of operational financial analysis. It retains the previous version’s emphasis on tool-calling, expert-verified financial reasoning, and correct calculations while extending the task to simulated investment banking and private equity workflows with incomplete requests, cross-source evidence, compliance constraints, and state changes.
The 200-task release is a representative frontier subset of the broader environment. Agents must resolve what a request refers to, select and interrogate the relevant data, perform valuation, diligence, risk, or transaction analysis, state assumptions, and complete any required update. The benchmark scores the workflow end to end, including whether the conclusion is correct, supported, safe, and reflected in the expected environment state.
At a glance
200
frontier tasks
93.5%
semantic-output tasks
69.5%
tasks with expected state changes
53.5%
tasks with 20K+ context
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
25.1%
|
37.4%
|
$3.46 | $17.29 |
| 2 | GLM 5.3 |
21.2%
|
42.3%
|
$0.88 | $4.42 |
| 3 | Fable 5.1 |
19.6%
|
34.2%
|
$5.83 | $29.16 |
| 4 | Kimi K3 |
17.3%
|
40.9%
|
$1.78 | $8.9 |
| 5 | Claude Opus 5 |
16.5%
|
33.5%
|
$3.2 | $16 |
| 6 | Qwen 3.8 Max |
14.1%
|
32.5%
|
$1.71 | $8.57 |
| 7 | GPT-6 Astra |
11.6%
|
19.3%
|
$10.6 | $53.02 |
| 8 | Muse Spark 1.3 |
10%
|
24%
|
$0.98 | $4.91 |
| 9 | DeepSeek V4 Pro |
9.5%
|
26.5%
|
$1.25 | $6.25 |
| 10 | Gemini Flash 3.8 |
5.4%
|
16%
|
$1.11 | $5.53 |
| 11 | Nemotron 3 Ultra 550B |
4.3%
|
13.6%
|
$0.42 | $2.09 |
Want to evaluate your model against SnorkelFinance 2.0? Talk to our team
Frontier performance
Methodology
Evaluator
Agents operate inside a Docker and OpenEnv environment with structured financial data, reference materials, MCP tools, stateful sessions, seeded failure behavior, and full trace logging. The setup requires evidence gathering and tool use within the environment.
scoring note
Scenario rubrics provide the primary score, with instance-level rubrics for selected complex tasks. Five calibration rollouts per task support difficulty measurement. Report task success with financial, state, safety, process, and efficiency diagnostics.
Behind the benchmark
Financial analysis in a deal team rarely begins with a complete question or a single source of truth. An agent may need to resolve the company or transaction, reconcile market and deal data, test assumptions, check restrictions, and communicate a conclusion that others can act on. SnorkelFinance 2.0 makes these dependencies part of the measured workflow, including the state changes required to complete the task.
Version 2.0 preserves V1’s expert-verified, tool-using financial reasoning and expands it from report-grounded QA into operational investment banking and private equity workflows. Experts define and review the workflows, data, rubrics, and verifiers. The evaluation now surfaces evidence traceability, analytical judgment, compliance-aware tool use, user interaction, and correct execution in a stateful deal environment.

