Agentic Coding
Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

Built with
Image
Image
Image
Image

At a glance

74

tasks in the v0.1 release

7

top-level domains in the taxonomy

31

subdomains spanning those 7 domains

35

rubric criteria every task must pass before merge

~34%

best model's pass rate
on v0.1, vs. ~5% for
the best open-weight model

Leaderboard

Rank Model Effort Agent Resolution Rate Release Date Tokens Cost
1 Opus 5 max mini-SWE-agent
42.7% ±1.6
2026-07-24 7.3B $5.8k
2 GPT-5.6 Sol max Codex
34.6% ±1.6
2026-07-09 5.8B $4.0k
3 Fable 5 max Claude Code
34.1% ±1.7
2026-06-09 3.6B $6.5k
4 GLM 5.3 max Claude Code
32.4% ±1.5
2026-08-14 5.6B $1.8k
5 Grok 4.6 high Grok Build
26.5% ±1.5
2026-08-12 2.9B $2.1k
6 Opus 4.8 max Claude Code
21.1% ±1.6
2026-05-28 5.2B $5.2k
7 GPT-5.6 Terra max Codex
20.8% ±1.4
2026-07-09 7.0B $2.5k
8 SWE-1.7 Lightning Devin
18.6% ±1.5
2026-07-08 3.6B $7.2k
9 Grok 4.5 xhigh Cursor CLI
15.7% ±1.5
2026-07-08 1.2B $766.02
10 Sonnet 5 max Claude Code
14.6% ±1.5
2026-06-30 17.9B $6.9k
11 GPT-5.6 Luna max Codex
14.3% ±1.3
2026-07-09 11.9B $1.6k
12 GLM 5.2 max Claude Code
4.6% ±1
2026-06-13 3.3B $3.4k

Key takeaways

Opus 5 (max effort) leads Terminal-Bench 3.0 at 42.7% ± 1.6% resolution, well ahead of GPT-5.6 Sol at 34.6% and Fable 5 at 34.1%. GLM 5.3 and Grok 4.6 follow at 32.4% and 26.5%, with the rest of the field spread down to 4.6%.

TAXONOMY

Seven domains, 31 subdomains

Every task is classified by the primary skill it exercises, not incidental tooling. The domain list is closed; subdomains grow as new tasks arrive.

science

Natural sciences & engineering

Biology

Chemistry
Physics
Earth
Robotics
Math
Linguistics

Software

General software engineering

Algorithms
Systems
Databases
Data engineering
Frontend
Languages

ML

Training, serving & eval

Training
Inference
Evaluation
Kernels

Operations

Business & financial reasoning

Finance
Logistics
Supply chain
Claims
Compliance
Marketing

Security

Offensive & defensive security

Cryptography
Reverse engineering
Forensics
AppSec

Hardware

Physical & digital hardware

CAD
RTL
Media

Creative & design work

Music
Design

Sample tasks

A selection from Snorkel AI's task contributions to Frontier-Bench's 74-task v0.1 release.

Review pipeline

Every task earns its place

Before a task counts toward Frontier-Bench, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.

01

Static checks

Canary, Dockerfile sanity, path & metadata validation — no API keys required.
02
Rubric review
An agent scores the task against 35 criteria: verifiable, solvable, difficult, anti-cheat robust.
03
Docker / Oracle / No-op
Environment builds; the reference solution passes; doing nothing fails.
04

Maintainer review

A maintainer reviews the task once all automated gates pass.
05
Agent & cheat trials
Multiple agents attempt the task; adversarial trials probe for reward hacking.
06
Hacker–fixer loop
An adversarial loop iteratively hardens the task against exploits before merge.
07
Merged
Counts toward Frontier-Bench and becomes eligible for the leaderboard.
of

How Terminal-Bench 3.0 compares

Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.

Benchmark

Task style

Domain span

Anti-cheat

Frontier-Bench

Community-submitted, expert-reviewed terminal tasks

7 domains / 31 subdomains

Rubric + cheat trials + hacker–fixer hardening loop

Terminal-Bench 2.1

89 fixed terminal tasks, continuously validated

Software, ML, security, data, science, sysadmin

Continuous validation, community-reported fixes

Senior SWE-Bench

Real PRs from 12 OSS repos, taste-graded

Software engineering only

Rubric + bloat + practice + relative-taste gates

OSWorld 2.0

Long-horizon computer-use workflows

7 professional domains, 31 self-hosted sites

Separate safety audit (8 checks)

Behind the benchmark

Terminal-Bench is used by virtually every frontier lab to measure whether agents can perform valuable work inside containerized terminal environments. Terminal-Bench 3.0 (formerly Frontier-Bench) is designed to track the frontier with a diverse, difficult, high quality set of tasks that evolve over time, pushing past the ceiling of 2.1.

Rather than a fixed release, Terminal-Bench 3.0 is being assembled through open community contribution. Anyone can propose and submit a task; every submission passes through an automated review pipeline — static checks, a 35-criteria implementation rubric, Docker/oracle/no-op validation, live agent trials, and adversarial "cheat" trials — before a maintainer signs off.

Snorkel AI contributes to Terminal-Bench 3.0 as both a task author and a data partner with additional support via the Open Benchmarks Grants program and Snorkel's Justin Bauer among the benchmark's reviewers.

Acknowledgments

Frontier-Bench is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Digital Work Index

An economically weighted measure of model performance across eight, frontier-difficulty benchmarks in software engineering, logical reasoning and writing, and professional environments.

Score
1
Image
Fable 5.1
23.62%
2
Image
GPT-6 Astra
23.03%
3
Image
Opus 5
21.75%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Fable 5.1
39.6%
3
Image
Opus 5
38.9%
Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Fable 5.1
14.5%
2
Image
Opus 5
14.0%
3
Image
GPT-6 Astra
12.4%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
GLM 5.3
30.4%
2
Image
DeepSeek V4 Pro
28.5%
3
Image
Kimi K3
27.9%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
GLM 5.3
11.9%
3
Image
Fable 5.1
11.3%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Fable 5.1
19.6%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Opus 5
14.9%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Scientific & Research Workflows

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
Fable 5.1
40.0%
2
Image
Opus 5
30.0%
3
Image
GPT-5.6 Sol
22.4%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By binary accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · xhigh
36.89%
3
Image
Opus 5 · high
33.33%
Software Engineering

Senior SWE-Bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
Image
GPT-6 Astra
34.2%
2
Image
Muse Spark 1.3
32.2%
3
Image
Opus 5
31.6%
Agentic Coding

Terminal-Bench 2.0

89 hard, human-verified terminal agent tasks in containerized environments, evaluated via task resolution rate.

Top models
1
Image
NexAU-AHE · GPT-5.5
84.7%
2
Image
LemonHarness · Multiple
84.5%
3
Image
Capy · GPT-5.5
83.1%

SnorkelWordle

A benchmark designed to evaluate linguistic reasoning and instruction-following capabilities in language models through the gameplay of Wordle.

Top Models
1
Image
GPT-5
94.0%
2
Image
Grok 4
93.0%
3
Image
o3
92.9%

SnorkelSpatial

A procedurally-generated benchmark for evaluating allocentric and egocentric spatial reasoning capabilities in LLMs.

Top Models
1
Image
GPT-5.4
99.0%
2
Image
Grok 4 Fast Reasoning
84.9%
3
Image
o3
76.7%

SnorkelGraph

A procedurally-generated and expert-verified benchmark for evaluating mathematical and spatial reasoning capabilities of LLMs through graph reasoning problems.

Top Models
1
Image
GPT-5.4
84.5%
2
Image
Grok 4 Fast Reasoning
75.0%
3
Image
o4-mini
75.0%

SnorkelSequences

A procedurally-generated and expert-verified benchmark for evaluating mathematical reasoning and compositional capabilities in LLMs.

Top Models
1
Image
GPT-5
77.6%
2
Image
GPT-5 Mini
77.6%
3
Image
GPT-5 Nano
72.0%
Enterprise Environments

SnorkelUnderwrite

An expert-verified frontier benchmark with multi-turn conversations, focused on agentic reasoning and tool use in commercial underwriting settings.

Top Models
1
Image
GPT-5.4
91.0%
2
Image
Claude Opus 4.1
86.3%
3
Image
GPT-5
83.3%
Enterprise Environments

SnorkelFinance

A benchmark of expert-verified financial QA created from financial reports for evaluating AI agents on tool-calling and reasoning capabilities.

Top Models
1
Image
GPT-5
81.0%
2
Image
o3
81.0%
3
Image
Gemini 3 Pro
80.3%
Enterprise Environments

Finance Reasoning

A benchmark co-created with Snorkel’s financial expert network, to test agents on financial reasoning questions through tool-calling and planning.

Top models
1
Image
Grok 4
53.1%
2
Image
GPT-5.4
52.0%
3
Image
Claude Sonnet 3.7
51.9%
Agentic Coding

Agentic Coding

A benchmark for evaluating AI models on complex, real-world coding tasks that require multi-step reasoning, tool use, and autonomous problem-solving.

Top Models
1
Image
Claude Opus 4.6
65.2%
2
Image
Claude Opus 4.5
58.0%
3
Image
Claude Sonnet 4.5
57.6%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

Top Systems (Agg. Reward)
1
Image
Sonnet 4.6 · ICL
+0.196
2
Image
GPT-5.4 · ICL
+0.189
3
Image
Sonnet 4.6 · Claude Code
+0.185
Agentic Coding

Terminal-Bench 2.1

Terminal agent evaluation led by Stanford University and Laude Institute. v2.1 fixes 28 tasks from 2.0 and introduces continuous validation.

Resolution Rate
1
Image
Codex CLI · GPT-5.5
83.4%
2
Image
Claude Code · Claude 5 Fable
83.1%
3
Image
Terminus 2 · Claude 5 Fable
80.4%
of

For models that need to be right. Not just good enough.