Model Insights Claude Opus 5 is the most consistent frontier performer in Snorkel’s current OBG suite, leading outright on Terminal-Bench 4.0, Agents’ Last Exam Overall/ALE-CLI/Last-Exam, and four of five TB-Science domains. Its one clear weak spot in this dataset is mathematical reasoning under TB-Science, where it drops to #3 behind Claude Fable 5 and GPT-5.6 Sol. Effort level matters more than