Image
Anthropic
·
RELEASED September 16, 2026

Model Insights

Claude Opus 5 is the most consistent frontier performer in Snorkel’s current OBG suite, leading outright on
Terminal-Bench 4.0, Agents’ Last Exam Overall/ALE-CLI/Last-Exam, and four of five TB-Science domains. Its one clear weak spot in this dataset is mathematical reasoning under TB-Science, where it drops to #3 behind Claude Fable 5 and GPT-5.6 Sol.

Effort level matters more than expected: on ALE Overall, High effort beats Max on both pass rate and cost (31.6% vs 30.9%, at $1,638 vs $2,302) — the most expensive setting is not the best one here. Use the effort selector above to see this across splits.

Snorkel benchmarks

Use the toggle to compare first-attempt reliability or success within five attempts.

Agentic Coding 2.0

Pass@1
38.9%
Pass@5
62%
Rank
#3/13
Rank
#1/13
Cost / trial
$2.91

Evaluation details
200 tasks · 979 graded trials

SWE-bench CLI+

Pass@1
14%
Pass@5
24.7%
Rank
#2/13
Rank
#2/13
Cost / trial
$23.73

Evaluation details
200 tasks · 990 graded trials

SnorkelWorkplace

Pass@1
14.9%
Pass@5
34.7%
Rank
#3/13
Rank
#2/13
Cost / trial
$15.96

Evaluation details
200 tasks · 970 graded trials

SnorkelFinance 2.0

Pass@1
16.5%
Pass@5
33.5%
Rank
#5/13
Rank
#5/13
Cost / trial
$3.2

Evaluation details
200 tasks · 982 graded trials

SnorkelUnderwrite 2.0

Pass@1
18.3%
Pass@5
34%
Rank
#9/13
Rank
#8/13
Cost / trial
$0.82

Evaluation details
200 tasks · 1000 graded trials

SnorkelManufacturing

Pass@1
2.9%
Pass@5
8.2%
Rank
#11/13
Rank
#11/13
Cost / trial
$3.34

Evaluation details
199 tasks · 995 graded trials

SnorkelRevOps

Pass@1
13%
Pass@5
33.8%
Rank
#6/13
Rank
#2/13
Cost / trial
$4.24

Evaluation details
197 tasks · 962 graded trials

SnorkelLegal

Pass@1
22.3%
Pass@5
43.7%
Rank
#8/13
Rank
#8/13
Cost / trial
$2.32

Evaluation details
199 tasks · 977 graded trials