Open Benchmarks Grants

Terminal-Bench-Science

Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

Built with
Image
Image
Image
Image

Overview

Terminal-Bench-Science evaluates agents on real scientific research workflows rather than software repository tasks. Version 0.1, the benchmark’s first public release, covers 70 tasks across five domain areas, each sourced from work a practicing researcher actually performed.

Headline finding: the strongest model evaluated, Claude Opus 5, resolves 30.0% of tasks in this release, well below the 50 to 80% range frontier models reach on general coding benchmarks. In this evaluation, authentic scientific workflows remain considerably harder for agents than general software engineering.

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

At a glance

70

tasks in the v0.1 release

5

domain areas covered

9

models evaluated

30.0%

top resolution rate (Claude Opus 5)
 

Leaderboard

Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
30% ±3.2
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
22.4% ±2.9
2026-07-09 8.4B $4.22K
3 Claude Fable 5 Claude Code max
21.4% ±2.8
2026-06-09 6.4B $14.18K
4 Claude Opus 4.8 Claude Code max
10.5% ±2.1
2026-05-28 6.8B $5.84K
5 GPT-5.6 Terra Codex max
8.6% ±1.9
2026-07-09 7.6B $1.98K
6 GLM 5.3 Claude Code max
8.1% ±1.9
2026-08-14 8.5B $2.73K
7 Kimi K3 Claude Code max
7.1% ±1.8
2026-07-16 3.2B $1.52K
8 Grok 4.6 Grok Build xhigh
7.1% ±1.8
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
3.3% ±1.2
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
29.8%
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
17.5%
2026-07-09 8.4B $4.22K
3 Claude Fable 5 Claude Code max
17.5%
2026-06-09 6.4B $14.18K
4 GLM 5.3 Claude Code max
10.5%
2026-08-14 8.5B $2.73K
5 Claude Opus 4.8 Claude Code max
7%
2026-05-28 6.8B $5.84K
6 GPT-5.6 Terra Codex max
7%
2026-07-09 7.6B $1.98K
7 Kimi K3 Claude Code max
7%
2026-07-16 3.2B $1.52K
8 Grok 4.6 Grok Build xhigh
7%
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
1.8%
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
45.8%
2026-07-24 7.3B $6.99K
2 Claude Fable 5 Claude Code max
25%
2026-06-09 6.4B $14.18K
3 GPT-5.6 Sol Codex max
20.8%
2026-07-09 8.4B $4.22K
4 Kimi K3 Claude Code max
12.5%
2026-07-16 3.2B $1.52K
5 GPT-5.6 Terra Codex max
8.3%
2026-07-09 7.6B $1.98K
6 Claude Opus 4.8 Claude Code max
4.2%
2026-05-28 6.8B $5.84K
7 GLM 5.3 Claude Code max
4.2%
2026-08-14 8.5B $2.73K
8 Grok 4.6 Grok Build xhigh
4.2%
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
0%
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
29.6%
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
14.8%
2026-07-09 8.4B $4.22K
3 Grok 4.6 Grok Build xhigh
14.8%
2026-08-12 3.7B $3.34K
4 Claude Fable 5 Claude Code max
11.1%
2026-06-09 6.4B $14.18K
5 Claude Opus 4.8 Claude Code max
7.4%
2026-05-28 6.8B $5.84K
6 GPT-5.6 Terra Codex max
7.4%
2026-07-09 7.6B $1.98K
7 GLM 5.3 Claude Code max
7.4%
2026-08-14 8.5B $2.73K
8 GPT-5.6 Luna Codex max
7.4%
2026-07-09 14.2B $383.20
9 Kimi K3 Claude Code max
3.7%
2026-07-16 3.2B $1.52K
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5 Claude Code max
33.3%
2026-06-09 6.4B $14.18K
2 GPT-5.6 Sol Codex max
31.4%
2026-07-09 8.4B $4.22K
3 Claude Opus 5 Claude Code max
25.5%
2026-07-24 7.3B $6.99K
4 GPT-5.6 Terra Codex max
9.8%
2026-07-09 7.6B $1.98K
5 Claude Opus 4.8 Claude Code max
7.8%
2026-05-28 6.8B $5.84K
6 GPT-5.6 Luna Codex max
7.8%
2026-07-09 14.2B $383.20
7 GLM 5.3 Claude Code max
5.9%
2026-08-14 8.5B $2.73K
8 Grok 4.6 Grok Build xhigh
5.9%
2026-08-12 3.7B $3.34K
9 Kimi K3 Claude Code max
2%
2026-07-16 3.2B $1.52K
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
27.5%
2026-07-24 7.3B $6.99K
2 GPT-5.6 Sol Codex max
23.5%
2026-07-09 8.4B $4.22K
3 Claude Opus 4.8 Claude Code max
21.6%
2026-05-28 6.8B $5.84K
4 Claude Fable 5 Claude Code max
17.6%
2026-06-09 6.4B $14.18K
5 Kimi K3 Claude Code max
11.8%
2026-07-16 3.2B $1.52K
6 GPT-5.6 Terra Codex max
9.8%
2026-07-09 7.6B $1.98K
7 GLM 5.3 Claude Code max
9.8%
2026-08-14 8.5B $2.73K
8 Grok 4.6 Grok Build xhigh
5.9%
2026-08-12 3.7B $3.34K
9 GPT-5.6 Luna Codex max
0%
2026-07-09 14.2B $383.20

Terminal-Bench-Science 0.1 Pareto Frontier

Loading chart data...

TAXONOMY

Domain areas

The v0.1 release spans five domain areas, unevenly weighted toward life and physical sciences. Resolution rates above are aggregated across all 70 tasks.

19

Life sciences

17

Physical sciences

8

Earth sciences

17

Mathematical sciences

9

Engineering sciences

Sample tasks

A sample of task titles from the v0.1 release. Browse the full contribution pipeline on the task dashboard →

Applied Mathematics / Physics

GRMHD primitive variable recovery

Recover physical (primitive) variables from conserved quantities in ideal general-relativistic magnetohydrodynamics under a strict double-precision arithmetic policy, graded against a hidden verifier corpus beyond the public sample cases.

Operations Research

Matrix-Free Computation of the Smallest Symplectic Eigenpairs
Implement a matrix-free solver for the invariant subspace of the smallest symplectic eigenvalues of a large implicit positive-definite operator, graded on symplectic feasibility, first-order stationarity, and certified-subspace accuracy under oracle, time, and memory limits.

Statistics / Ocean Sciences

Honest posterior intervals for a mixed-layer ocean model

Produce honest Bayesian posterior intervals from correlated, instrument-biased ocean temperature-profile residuals, where a naive independent-Gaussian likelihood passes standard convergence diagnostics while returning intervals several times too narrow.

Biology

Add task: crispr-mhcii-screen
Analyze deposited CRISPR-Cas9 pooled-screen count matrices from unequal-depth melanoma MHC-II FACS replicates to produce a defensible gene hit list, with no processed reference hit list provided.

Biology

Genome-to-phenotype prediction under multi-cohort lineage shift
Identify genuine causal genome-to-phenotype markers across six family-disjoint cohorts where decoy modules are statistically identical to genuine ones in labeled data but independently rewired per deployment cohort, defeating ordinary cross-validation.

Physics

Self-Consistent Magnetic Phase Boundaries in Correlated Moiré Models
Solve for self-consistent magnetic phase boundaries in spin-orbit-coupled moiré Hamiltonians, requiring simultaneous convergence of coupled impurity self-energies and structural relaxation with general spin mixing.

Chemistry

Hydration free energy calculation of a solute model
Compute a solute's hydration free energy despite a slow conformational degree of freedom orthogonal to the alchemical variable, where a routine calculation gets trapped in one metastable state and reports a confidently wrong, deceptively low-uncertainty result.

Astronomy

Fast Radio Burst Reconstruction
Detect and reconstruct a fast radio burst from noisy time-series data under strict timing and precision requirements, verified against dispersion-measure and pulse-width tolerances.
Review pipeline

How tasks are graded

Every task in Terminal-Bench-Science goes through a review pipeline before it is eligible to be merged into a release, and every evaluation run separates the agent’s working environment from the environment that grades it.

01

Proposal

A domain expert or contributor proposes a task drawn from real research work.
02

LLM Judge

An automated judge screens the proposal for basic feasibility and completeness.
03

Reviewer

A first-pass human reviewer checks instruction clarity and task scope.
04

Senior Reviewer

A senior reviewer checks instruction and verifier alignment, task realism, and under or over specification risk.
05

Pull Request

The task is opened as a pull request against the public task repository.
06

Static Checks

Automated checks validate task structure, environment definitions, and verifier code.

07

Agent Judge

Frontier, oracle, and cheating-agent trial runs confirm the task is solvable, gradeable, and resistant to shortcuts.

08

Agent Trials

Additional agent trial runs stress test the task before it is accepted.

09

Merge

The task is merged into the public task set and becomes eligible for a future release.
of

Acknowledgments

Terminal-Bench-Science is a collaboration between Stanford University, the Laude Institute, and domain experts across scientific disciplines and research institutions worldwide, led by Steven Dillmann (Stanford University), Sanmi Koyejo (Stanford University), and Ludwig Schmidt (Stanford University, Anthropic)

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

New
Open Benchmarks Grants

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
Image
Opus 5 (max)
51.8%
2
Image
Fable 5 (max)
44.5%
3
Image
GLM 5.3 (max)
41.8%
Open Benchmarks Grants

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
GPT-5.6 Sol (Codex)
34.4%
2
Image
Fable 5 (Claude Code)
33.8%
3
Image
Opus 4.8 (Claude Code)
21.1%
Open Benchmarks Grants

OSWorld 2.0

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By binary accuracy (300 steps)
1
Image
gpt-5-5 · xhigh
13%
2
Image
Claude Opus 4.7 · max
13%
3
Image
Claude Sonnet 4.6 · medium
8.3%
New

Senior SWE-Bench

A benchmark for evaluating coding agents on senior-level engineering work: building features from realistic instructions, investigating bugs that require runtime investigation, and shipping code that aligns to existing codebase conventions.

Tasteful Solve Rate
1
Image
Claude Fable 5
27.9%
2
Image
Claude Opus 4.8
25.0%
3
Image
Claude Sonnet 5
17.4%
Open Benchmarks Grants

Agents’ Last Exam

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By Binary Accuracy
1
Image
Codex · GPT-5.5
24%
2
Image
ALE Claw · GPT-5.5
23%
3
Image
Claude Code · Claude-Fable-5
22%

Agentic Coding

A benchmark for evaluating AI models on complex, real-world coding tasks that require multi-step reasoning, tool use, and autonomous problem-solving.

Top Models
1
Image
Claude Opus 4.6
65.2%
2
Image
Claude Opus 4.5
58.0%
3
Image
Claude Sonnet 4.5
57.6%
Open Benchmarks Grants

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Open Benchmarks Grants

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
ICL · Claude Sonnet 4.6
+0.223
2
Image
ICL · GPT-5.4
+0.201
3
Image
Claude Code · Claude Sonnet 4.6
+0.190
of

For models that need to be right. Not just good enough.