Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for terminal agents. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

Built with
Image
Image
Image
Image

Overview

Terminal-Bench is used by virtually every frontier lab to measure whether agents can perform valuable work inside containerized terminal environments. Terminal-Bench 3.0 (formerly Frontier-Bench) is designed to track the frontier with a diverse, difficult, high quality set of tasks that evolve over time, pushing past the ceiling of 2.1, where top agents already clear 75–84%.

Rather than a fixed release, Terminal-Bench 3.0 is being assembled through open community contribution. Anyone can propose and submit a task; every submission passes through an automated review pipeline — static checks, a 35-criteria implementation rubric, Docker/oracle/no-op validation, live agent trials, and adversarial "cheat" trials — before a maintainer signs off.

Snorkel AI contributes to Terminal-Bench 3.0 as both a task author and a data partner with additional support via the Open Benchmarks Grants program and Snorkel's Justin Bauer among the benchmark's reviewers.

At a glance

74

tasks in the v0.1 release

7

top-level domains in the taxonomy

31

subdomains spanning those 7 domains

35

rubric criteria every task must pass before merge

~34%

best model's pass rate on v0.1, vs. ~5% for the best open-weight model

Leaderboard

Rank Model Agent Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 (max) mini-SWE-agent
42.7% ±1.6
2026-07-24 7.3B $5.8k
2 GPT-5.6 Sol (max) Codex
34.6% ±1.6
2026-07-09 5.8B $4.0k
3 Claude Fable 5 (max) Claude Code
34.1% ±1.7
2026-06-09 3.6B $6.5k
4 GLM-5.3 (max) Claude Code
32.4% ±1.5
2026-08-14 5.6B $1.8k
5 Grok 4.6 (high) Grok Build
26.5% ±1.5
2026-08-12 2.9B $2.1k
6 Claude Opus 4.8 (max) Claude Code
21.1% ±1.6
2026-05-28 5.2B $5.2k
7 GPT-5.6 Terra (max) Codex
20.8% ±1.4
2026-07-09 7.0B $2.5k
8 SWE-1.7 Lightning Devin
18.6% ±1.5
2026-07-08 3.6B $7.2k
9 Grok 4.5 (xhigh) Cursor CLI
15.7% ±1.5
2026-07-08 1.2B $766.02
10 Claude Sonnet 5 (max) Claude Code
14.6% ±1.5
2026-06-30 17.9B $6.9k
11 GPT-5.6 Luna (max) Codex
14.3% ±1.3
2026-07-09 11.9B $1.6k
12 GLM-5.2 (max) Claude Code
4.6% ±1
2026-06-13 3.3B $3.4k

TAXONOMY

Seven domains, 31 subdomains

Every task is classified by the primary skill it exercises, not incidental tooling. The domain list is closed; subdomains grow as new tasks arrive.

science

Natural sciences & engineering

Biology

Chemistry
Physics
Earth
Robotics
Math
Linguistics

Software

General software engineering

Algorithms
Systems
Databases
Data engineering
Frontend
Languages

ML

Training, serving & eval

Training
Inference
Evaluation
Kernels

Operations

Business & financial reasoning

Finance
Logistics
Supply chain
Claims
Compliance
Marketing

Security

Offensive & defensive security

Cryptography
Reverse engineering
Forensics
AppSec

Hardware

Physical & digital hardware

CAD
RTL
Media

Creative & design work

Music
Design

Sample tasks

A selection from Snorkel AI's task contributions to Frontier-Bench's 74-task v0.1 release.

Review pipeline

Every task earns its place

Before a task counts toward Frontier-Bench, it passes through an automated + human review pipeline designed to catch broken tasks and reward-hacking before they ever reach the leaderboard.

01

Static checks

Canary, Dockerfile sanity, path & metadata validation — no API keys required.
02
Rubric review
An agent scores the task against 35 criteria: verifiable, solvable, difficult, anti-cheat robust.
03
Docker / Oracle / No-op
Environment builds; the reference solution passes; doing nothing fails.
04

Maintainer review

A maintainer reviews the task once all automated gates pass.
05
Agent & cheat trials
Multiple agents attempt the task; adversarial trials probe for reward hacking.
06
Hacker–fixer loop
An adversarial loop iteratively hardens the task against exploits before merge.
07
Merged
Counts toward Frontier-Bench and becomes eligible for the leaderboard.
of

How Frontier-Bench compares

Several recent benchmarks make progress on behavioral testing and instruction realism. The following table provides a brief comparison.

Benchmark

Task style

Domain span

Anti-cheat

Frontier-Bench

Community-submitted, expert-reviewed terminal tasks

7 domains / 31 subdomains

Rubric + cheat trials + hacker–fixer hardening loop

Terminal-Bench 2.1

89 fixed terminal tasks, continuously validated

Software, ML, security, data, science, sysadmin

Continuous validation, community-reported fixes

Senior SWE-Bench

Real PRs from 12 OSS repos, taste-graded

Software engineering only

Rubric + bloat + practice + relative-taste gates

OSWorld 2.0

Long-horizon computer-use workflows

7 professional domains, 31 self-hosted sites

Separate safety audit (8 checks)

Acknowledgments

Frontier-Bench is hosted by Harbor and the Laude Institute, led by Ryan Marten, Alex Shaw, Andy Konwinski, and Ludwig Schmidt, with data and research contributions from Snorkel AI led by Justin Bauer and Vincent Sunn Chen, and additional support through Snorkel's Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Fable 5.1
39.6%
3
Image
Opus 5
38.9%
Software Engineering

SWE-bench CLI

Tests whether an AI coding agent can diagnose and deliver a validated, multi-file change in a real open-source repository, navigating code, tests, dependencies, and tooling through the command line.

By pass@1
1
Image
Fable 5.1
14.5%
2
Image
Opus 5
14.0%
3
Image
GPT-6 Astra
12.4%
Enterprise Environments

SnorkelUnderwrite 2.0

Measures an agent’s ability to turn incomplete, distributed insurance evidence into an auditable underwriting decision while respecting authority and policy constraints.

By pass@1
1
Image
GLM 5.3
30.4%
2
Image
DeepSeek V4 Pro
28.5%
3
Image
Kimi K3
27.9%
Enterprise Environments

SnorkelManufacturing

Measures how well AI agents can turn fragmented plant-floor, engineering, and supplier evidence into safe, technically defensible decisions and actions.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
GLM 5.3
11.9%
3
Image
Fable 5.1
11.3%
Enterprise Environments

SnorkelRevOps

Tests whether AI agents can reconcile revenue systems, enforce hard commercial controls, and carry a decision through to the correct business state.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Scores how agents gather evidence, perform financial analysis, follow compliance constraints, and complete required state updates in a simulated environment.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Fable 5.1
19.6%
Enterprise Environments

SnorkelLegal

Evaluates AI agents on whether they can advance a legal matter while preserving the chain from controlling evidence to procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Muse Spark 1.3
17.7%
2
Image
Grok 4.6
17.1%
3
Image
Fable 5.1
15.8%
Agentic Coding

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Scientific & Research Workflows

Terminal-Bench-Science

A benchmark for evaluating AI agents on workflows from researchers’ own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

By Resolution Rate
1
Image
Fable 5.1
40.0%
2
Image
Opus 5
30.0%
3
Image
GPT-5.6 Sol
22.4%
Computer Use

OSWorld 2.0

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By binary accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · xhigh
36.89%
3
Image
Opus 5 · high
33.33%
Software Engineering

Senior SWE-Bench

Evaluating coding agents on senior-level engineering work.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
Image
GPT-6 Astra
34.2
2
Image
Muse Spark 1.3
32.2%
3
Image
Opus 5
31.6%
Software Engineering

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
Sonnet 4.6 · ICL
+0.196
2
Image
GPT-5.4 · ICL
+0.189
3
Image
Sonnet 4.6 · Claude Code
+0.185
of

For models that need to be right. Not just good enough.