Scientific & Research Workflows
Open Benchmarks Grants

Terminal-Bench-Science

Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or vendors, set the bar for scientific capability in AI.

Built with
Image
Image
Image
Image

Overview

Terminal-Bench-Science evaluates agents on real scientific research workflows rather than software repository tasks. Version 0.1, the benchmark’s first public release, covers 70 tasks across five domain areas, each sourced from work a practicing researcher actually performed.

Headline finding: the strongest model evaluated, Claude Opus 5, resolves 30.0% of tasks in this release, well below the 50 to 80% range frontier models reach on general coding benchmarks. In this evaluation, authentic scientific workflows remain considerably harder for agents than general software engineering.

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

At a glance

70

tasks in the v0.1 release

5

domain areas covered

30.0%

top resolution rate (Claude Opus 5)
 

Leaderboard

Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5.1 Claude Code max
40% ±3.4
2026-09-01 3.8B $6.25K
2 Claude Opus 5 Claude Code max
30% ±3.2
2026-07-24 7.3B $6.99K
3 GPT-5.6 Sol Codex max
22.4% ±2.9
2026-07-09 8.4B $4.22K
4 Claude Fable 5 Claude Code max
21.4% ±2.8
2026-06-09 6.4B $14.18K
5 DeepSeek V4.1 Flash Codex max
15.7% ±2.5
2026-09-10 14.9B $385.93
6 Gemini 3.8 Flash mini-SWE-agent high
12.4% ±2.3
2026-09-02 10.5B $1.12K
7 Claude Opus 4.8 Claude Code max
10.5% ±2.1
2026-05-28 6.8B $5.84K
8 GPT-5.6 Terra Codex max
8.6% ±1.9
2026-07-09 7.6B $1.98K
9 GLM 5.3 Claude Code max
8.1% ±1.9
2026-08-14 8.5B $2.73K
10 Kimi K3 Claude Code max
7.1% ±1.8
2026-07-16 3.2B $1.52K
11 Grok 4.6 Grok Build xhigh
7.1% ±1.8
2026-08-12 3.7B $3.34K
12 Gemini 3.7 Flash mini-SWE-agent high
5.7% ±1.6
2026-08-13 12.0B $1.30K
13 DeepSeek V4 Pro Codex max
3.8% ±1.3
2026-08-13 11.2B $1,120.15
14 GPT-5.6 Luna Codex max
3.3% ±1.2
2026-07-09 14.2B $383.20
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5.1 Claude Code max
38.6% ±6.4
2026-09-01 751.4M $1.11K
2 Claude Opus 5 Claude Code max
29.8% ±6.1
2026-07-24 1.3B $1.32K
3 GPT-5.6 Sol Codex max
17.5% ±5
2026-07-09 911.2M $513.19
4 Claude Fable 5 Claude Code max
17.5% ±5
2026-06-09 864.3M $1.96K
5 Gemini 3.8 Flash mini-SWE-agent high
14% ±4.6
2026-09-02 947.5M $129.49
6 GLM 5.3 Claude Code max
10.5% ±4.1
2026-08-14 1.5B $1.31K
7 DeepSeek V4.1 Flash Codex max
10.5% ±4.1
2026-09-10 2.5B $64.23
8 Gemini 3.7 Flash mini-SWE-agent high
8.8% ±3.7
2026-08-13 704.8M $100.54
9 Grok 4.6 Grok Build xhigh
7% ±3.4
2026-08-12 367.5M $281.65
10 GPT-5.6 Terra Codex max
7% ±3.4
2026-07-09 1B $297.96
11 Kimi K3 Claude Code max
7% ±3.4
2026-07-16 612.8M $591.34
12 Claude Opus 4.8 Claude Code max
7% ±3.4
2026-05-28 780.5M $806.35
13 GPT-5.6 Luna Codex max
1.8% ±1.7
2026-07-09 2.9B $83.51
14 DeepSeek V4 Pro Codex max
0%
2026-08-13 2.1B $107.43
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5.1 Claude Code max
37.3% ±6.8
2026-09-01 1B $1.82K
2 Claude Opus 5 Claude Code max
27.5% ±6.2
2026-07-24 2.1B $1.85K
3 DeepSeek V4.1 Flash Codex max
27.5% ±6.2
2026-09-10 4.6B $118.82
4 Gemini 3.8 Flash mini-SWE-agent high
23.5% ±5.9
2026-09-02 2.8B $288.67
5 GPT-5.6 Sol Codex max
23.5% ±5.9
2026-07-09 2.4B $1.14K
6 Claude Opus 4.8 Claude Code max
21.6% ±5.8
2026-05-28 2B $1.6K
7 Claude Fable 5 Claude Code max
17.6% ±5.3
2026-06-09 1.5B $3.58K
8 Kimi K3 Claude Code max
11.8% ±4.5
2026-07-16 799.3M $536.23
9 GPT-5.6 Terra Codex max
9.8% ±4.2
2026-07-09 3.3B $791.88
10 GLM 5.3 Claude Code max
9.8% ±4.2
2026-08-14 2.3B $1.73K
11 Gemini 3.7 Flash mini-SWE-agent high
7.8% ±3.8
2026-08-13 4.7B $490.74
12 Grok 4.6 Grok Build xhigh
5.9% ±3.3
2026-08-12 646.9M $539.59
13 DeepSeek V4 Pro Codex max
3.9% ±2.7
2026-08-13 2.8B $140.45
14 GPT-5.6 Luna Codex max
0%
2026-07-09 4.8B $116.84
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Opus 5 Claude Code max
45.8% ±10.2
2026-07-24 897.5M $1.15K
2 Claude Fable 5.1 Claude Code max
37.5% ±9.9
2026-09-01 431.2M $746.48
3 Claude Fable 5 Claude Code max
25% ±8.8
2026-06-09 602.6M $1.62K
4 GPT-5.6 Sol Codex max
20.8% ±8.3
2026-07-09 1.2B $584.56
5 Gemini 3.8 Flash mini-SWE-agent high
12.5% ±6.8
2026-09-02 1.1B $117.58
6 Kimi K3 Claude Code max
12.5% ±6.8
2026-07-16 380.1M $342.60
7 GPT-5.6 Terra Codex max
8.3% ±5.6
2026-07-09 1.2B $302.41
8 DeepSeek V4.1 Flash Codex max
8.3% ±5.6
2026-09-10 1.8B $39.17
9 Grok 4.6 Grok Build xhigh
4.2% ±4.1
2026-08-12 244.9M $217.33
10 Claude Opus 4.8 Claude Code max
4.2% ±4.1
2026-05-28 753.7M $694.26
11 GLM 5.3 Claude Code max
4.2% ±4.1
2026-08-14 1.6B $1.4K
12 GPT-5.6 Luna Codex max
0%
2026-07-09 1.6B $41.41
13 Gemini 3.7 Flash mini-SWE-agent high
0%
2026-08-13 503.6M $63.51
14 DeepSeek V4 Pro Codex max
0%
2026-08-13 1.5B $71.82
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5.1 Claude Code max
48.1% ±9.6
2026-09-01 633.4M $976.65
2 Claude Opus 5 Claude Code max
29.6% ±8.8
2026-07-24 1.2B $1.04K
3 DeepSeek V4.1 Flash Codex max
18.5% ±7.5
2026-09-10 3B $69.75
4 Grok 4.6 Grok Build xhigh
14.8% ±6.8
2026-08-12 276.7M $231.57
5 GPT-5.6 Sol Codex max
14.8% ±6.8
2026-07-09 569.2M $315.06
6 Gemini 3.8 Flash mini-SWE-agent high
11.1% ±6
2026-09-02 520.2M $70.26
7 DeepSeek V4 Pro Codex max
11.1% ±6
2026-08-13 2.2B $108.40
8 Claude Fable 5 Claude Code max
11.1% ±6
2026-06-09 1.3B $2.39K
9 GPT-5.6 Luna Codex max
7.4% ±5
2026-07-09 1.2B $35.82
10 Gemini 3.7 Flash mini-SWE-agent high
7.4% ±5
2026-08-13 377.2M $56.41
11 GPT-5.6 Terra Codex max
7.4% ±5
2026-07-09 464.1M $141.00
12 GLM 5.3 Claude Code max
7.4% ±5
2026-08-14 992.1M $788.27
13 Claude Opus 4.8 Claude Code max
7.4% ±5
2026-05-28 1.1B $831.49
14 Kimi K3 Claude Code max
3.7% ±3.6
2026-07-16 536.3M $354.73
Rank Model Agent Effort Resolution Rate Release Date Tokens Cost
1 Claude Fable 5.1 Claude Code max
41.2% ±6.9
2026-09-01 891.4M $1.6K
2 Claude Fable 5 Claude Code max
33.3% ±6.6
2026-06-09 2.1B $4.63K
3 GPT-5.6 Sol Codex max
31.4% ±6.5
2026-07-09 3.3B $1.66K
4 Claude Opus 5 Claude Code max
25.5% ±6.1
2026-07-24 1.7B $1.63K
5 DeepSeek V4.1 Flash Codex max
11.8% ±4.5
2026-09-10 2.9B $93.97
6 GPT-5.6 Terra Codex max
9.8% ±4.2
2026-07-09 1.6B $450.66
7 GPT-5.6 Luna Codex max
7.8% ±3.8
2026-07-09 3.7B $105.62
8 Claude Opus 4.8 Claude Code max
7.8% ±3.8
2026-05-28 2.2B $1.9K
9 DeepSeek V4 Pro Codex max
5.9% ±3.3
2026-08-13 2.7B $146.13
10 GLM 5.3 Claude Code max
5.9% ±3.3
2026-08-14 2.1B $1.57K
11 Grok 4.6 Grok Build xhigh
5.9% ±3.3
2026-08-12 2.2B $2.07K
12 Gemini 3.7 Flash mini-SWE-agent high
2% ±1.9
2026-08-13 5.7B $585.39
13 Kimi K3 Claude Code max
2% ±1.9
2026-07-16 863.2M $714.76
14 Gemini 3.8 Flash mini-SWE-agent high
0%
2026-09-02 5.2B $510.98

Terminal-Bench-Science 0.1 Pareto Frontier

Loading chart data...

TAXONOMY

Domain areas

The v0.1 release spans five domain areas, unevenly weighted toward life and physical sciences. Resolution rates above are aggregated across all 70 tasks.

19

Life sciences

17

Physical sciences

8

Earth sciences

17

Mathematical sciences

9

Engineering sciences

Sample tasks

A sample of task titles from the v0.1 release. Browse the full contribution pipeline on the task dashboard →

Applied Mathematics / Physics

GRMHD primitive variable recovery

Recover physical (primitive) variables from conserved quantities in ideal general-relativistic magnetohydrodynamics under a strict double-precision arithmetic policy, graded against a hidden verifier corpus beyond the public sample cases.

Operations Research

Matrix-Free Computation of the Smallest Symplectic Eigenpairs
Implement a matrix-free solver for the invariant subspace of the smallest symplectic eigenvalues of a large implicit positive-definite operator, graded on symplectic feasibility, first-order stationarity, and certified-subspace accuracy under oracle, time, and memory limits.

Statistics / Ocean Sciences

Honest posterior intervals for a mixed-layer ocean model

Produce honest Bayesian posterior intervals from correlated, instrument-biased ocean temperature-profile residuals, where a naive independent-Gaussian likelihood passes standard convergence diagnostics while returning intervals several times too narrow.

Biology

Add task: crispr-mhcii-screen
Analyze deposited CRISPR-Cas9 pooled-screen count matrices from unequal-depth melanoma MHC-II FACS replicates to produce a defensible gene hit list, with no processed reference hit list provided.

Biology

Genome-to-phenotype prediction under multi-cohort lineage shift
Identify genuine causal genome-to-phenotype markers across six family-disjoint cohorts where decoy modules are statistically identical to genuine ones in labeled data but independently rewired per deployment cohort, defeating ordinary cross-validation.

Physics

Self-Consistent Magnetic Phase Boundaries in Correlated Moiré Models
Solve for self-consistent magnetic phase boundaries in spin-orbit-coupled moiré Hamiltonians, requiring simultaneous convergence of coupled impurity self-energies and structural relaxation with general spin mixing.

Chemistry

Hydration free energy calculation of a solute model
Compute a solute's hydration free energy despite a slow conformational degree of freedom orthogonal to the alchemical variable, where a routine calculation gets trapped in one metastable state and reports a confidently wrong, deceptively low-uncertainty result.

Astronomy

Fast Radio Burst Reconstruction
Detect and reconstruct a fast radio burst from noisy time-series data under strict timing and precision requirements, verified against dispersion-measure and pulse-width tolerances.
Review pipeline

How tasks are graded

Every task in Terminal-Bench-Science goes through a review pipeline before it is eligible to be merged into a release, and every evaluation run separates the agent’s working environment from the environment that grades it.

01

Proposal

A domain expert or contributor proposes a task drawn from real research work.
02

LLM Judge

An automated judge screens the proposal for basic feasibility and completeness.
03

Reviewer

A first-pass human reviewer checks instruction clarity and task scope.
04

Senior Reviewer

A senior reviewer checks instruction and verifier alignment, task realism, and under or over specification risk.
05

Pull Request

The task is opened as a pull request against the public task repository.
06

Static Checks

Automated checks validate task structure, environment definitions, and verifier code.

07

Agent Judge

Frontier, oracle, and cheating-agent trial runs confirm the task is solvable, gradeable, and resistant to shortcuts.

08

Agent Trials

Additional agent trial runs stress test the task before it is accepted.

09

Merge

The task is merged into the public task set and becomes eligible for a future release.
of

Acknowledgments

Terminal-Bench-Science is a collaboration between Stanford University, the Laude Institute, and domain experts across scientific disciplines and research institutions worldwide, led by Steven Dillmann (Stanford University), Sanmi Koyejo (Stanford University), and Ludwig Schmidt (Stanford University, Anthropic)

Snorkel supports Terminal-Bench Science through the Open Benchmarks Grants program.

FAQs

Get notified when we launch a new benchmark

Share this benchmark

More benchmarks

Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Fable 5.1
39.6%
3
Image
Opus 5
38.9%
Software Engineering

SWE-bench CLI

Tests whether an AI coding agent can diagnose and deliver a validated, multi-file change in a real open-source repository, navigating code, tests, dependencies, and tooling through the command line.

By pass@1
1
Image
Fable 5.1
14.5%
2
Image
Opus 5
14.0%
3
Image
GPT-6 Astra
12.4%
Enterprise Environments

SnorkelUnderwrite 2.0

Measures an agent’s ability to turn incomplete, distributed insurance evidence into an auditable underwriting decision while respecting authority and policy constraints.

By pass@1
1
Image
GLM 5.3
30.4%
2
Image
DeepSeek V4 Pro
28.5%
3
Image
Kimi K3
27.9%
Enterprise Environments

SnorkelManufacturing

Measures how well AI agents can turn fragmented plant-floor, engineering, and supplier evidence into safe, technically defensible decisions and actions.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
GLM 5.3
11.9%
3
Image
Fable 5.1
11.3%
Enterprise Environments

SnorkelRevOps

Tests whether AI agents can reconcile revenue systems, enforce hard commercial controls, and carry a decision through to the correct business state.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Scores how agents gather evidence, perform financial analysis, follow compliance constraints, and complete required state updates in a simulated environment.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Fable 5.1
19.6%
Enterprise Environments

SnorkelLegal

Evaluates AI agents on whether they can advance a legal matter while preserving the chain from controlling evidence to procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Muse Spark 1.3
17.7%
2
Image
Grok 4.6
17.1%
3
Image
Fable 5.1
15.8%
Agentic Coding

Terminal-Bench 4.0

A benchmark to measure and evolve with the frontier of agent work: real terminal environments, real software engineering tasks, and a rolling task set that is revised as agents catch up to it.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Agentic Coding

Terminal-Bench 3.0

The next frontier benchmark for agent work. A harder, more domain-diverse successor to Terminal-Bench 2.1 — built in the open, task by task, under continuous adversarial review.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Computer Use

OSWorld 2.0

Long-horizon professional workflows with verifiable outcomes across 55 sub-industries. 147 public tasks of a 1,500+ task corpus, sourced and validated by 300+ industry experts.

By binary accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · xhigh
36.89%
3
Image
Opus 5 · high
33.33%
Software Engineering

Senior SWE-Bench

Evaluating coding agents on senior-level engineering work.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
Image
GPT-6 Astra
34.2
2
Image
Muse Spark 1.3
32.2%
3
Image
Opus 5
31.6%
Software Engineering

SlopCode Bench

Measures code quality degradation in AI-assisted codebases. Tracks checkpoint solve rates, erosion (code bloat), and verbosity under realistic repo conditions.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Evaluates whether AI systems improve from prior experience across sequential, stateful tasks, measuring real in-context learning, not just raw capability.

Top Systems (Agg. Reward)
1
Image
Sonnet 4.6 · ICL
+0.196
2
Image
GPT-5.4 · ICL
+0.189
3
Image
Sonnet 4.6 · Claude Code
+0.185
of

For models that need to be right. Not just good enough.