Model Insights
Claude Opus 5 is the most consistent frontier performer in Snorkel’s current OBG suite, leading outright on Terminal-Bench 4.0, Agents’ Last Exam Overall/ALE-CLI/Last-Exam, and four of five TB-Science domains. Its one clear weak spot in this dataset is mathematical reasoning under TB-Science, where it drops to #3 behind Claude Fable 5 and GPT-5.6 Sol.
Effort level matters more than expected: on ALE Overall, High effort beats Max on both pass rate and cost (31.6% vs 30.9%, at $1,638 vs $2,302) — the most expensive setting is not the best one here. Use the effort selector above to see this across splits.
Snorkel benchmarks
Agentic Coding 2.0
Evaluation details
200 tasks · 979 graded trials
SWE-bench CLI+
Evaluation details
200 tasks · 990 graded trials
SnorkelWorkplace
Evaluation details
200 tasks · 970 graded trials
SnorkelFinance 2.0
Evaluation details
200 tasks · 982 graded trials
SnorkelUnderwrite 2.0
Evaluation details
200 tasks · 1000 graded trials
SnorkelManufacturing
Evaluation details
199 tasks · 995 graded trials
SnorkelRevOps
Evaluation details
197 tasks · 962 graded trials
SnorkelLegal
Evaluation details
199 tasks · 977 graded trials