SnorkelLegal
Evaluates AI agents on whether they can advance a legal matter while preserving the chain from controlling evidence to procedure, authority, and action.
overview
SnorkelLegal evaluates agents within a simulated mid-market law firm, where every request is tied to a matter, a procedural stage, and a role with defined authority.
The release covers representative work across matter intake, conflicts, docketing, litigation holds, discovery, privilege, settlement authority, transactional exposure, and regulatory investigations. Agents must locate the correct matter and parties, find the controlling document or policy, distinguish current evidence from superseded material, and return a bounded legal work product or make a governed matter update.
The benchmark evaluates the intersection of long-context reasoning, matter-specific rules, and procedural discipline.
At a glance
200
frontier tasks
15
simulated personas represented
50%
semantic-output tasks
23.5%
tasks requiring state changes
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
40.7%
|
64.6%
|
$2.22 | $11.12 |
| 2 | GLM 5.3 |
37.8%
|
63.6%
|
$1.37 | $6.83 |
| 3 | Muse Spark 1.3 |
36.5%
|
55.5%
|
$1.45 | $7.25 |
| 4 | Kimi K3 |
36.1%
|
65.3%
|
$1.94 | $9.69 |
| 5 | Fable 5.1 |
35.9%
|
60.5%
|
$5.26 | $26.31 |
| 6 | DeepSeek V4 Pro |
30.4%
|
63.5%
|
$1.37 | $6.84 |
| 7 | Claude Opus 5 |
26.4%
|
50%
|
$4.14 | $20.69 |
| 8 | Qwen 3.8 Max |
25.6%
|
53.5%
|
$2.34 | $11.69 |
| 9 | Gemini Flash 3.8 |
25.3%
|
42%
|
$1.23 | $6.16 |
| 10 | GPT-6 Astra |
20.9%
|
35.1%
|
$11.61 | $58.04 |
Want to evaluate your model against SnorkelLegal? Talk to our team
Frontier performance
Methodology
Evaluator
Each task runs inside a reproducible Docker/OpenEnv workspace with matter records, document evidence, role-gated MCP tools, and a simulated requester. The session preserves matter state and captures the complete trajectory, including retrieval, authority checks, user interaction, and governed writes. Seeded randomness supports repeatable failure-mode testing.
scoring note
Fluent legal prose is not enough. A task is incorrect when the agent relies on superseded evidence, skips a required authority or conflict check, makes an impermissible write, or gives a semantically wrong answer that merely sounds plausible. The score combines outcome gates with process and efficiency diagnostics while allowing valid alternative tool sequences.
Behind the benchmark
Legal work is not just retrieval plus prose. The operative rule may be firm-private or jurisdiction-specific. A matter file may place governing text next to superseded drafts and withdrawn analyses. Authority and confidentiality can determine whether an otherwise correct action is allowed.
SnorkelLegal turns those constraints into observable tests of grounding, sequencing, and restraint. The subset includes semantic legal outputs, long-context tasks, and governed matter updates. It asks a practical question for model developers: can an agent move a matter forward without inventing certainty, using the wrong version, or bypassing a control?
Tasks and rubrics are authored and reviewed by legal-domain experts.

