We build the data that pushes the frontier

Snorkel helps AI labs develop specialized training data and environments that set their models and agents apart.

Proud to partner with top frontier AI and research teams
Image
Image
Image
Image
Image
Image
Image
Image
Image
Image
Image

Frontier models break at the edges. We build for that.

Most data pipelines are built for volume, not difficulty. Frontier models fail on distributional gaps in specialized domains, benchmark blind spots, and tasks where correctness is hard to define. Snorkel is built specifically for these problems.
Frontier models inforgraphics

Founded out of Stanford AI Lab, we've been shaping and benchmarking frontier AI for nearly a decade.

Leaderboards
Software Engineering

Senior SWE-Bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
4
Image
GPT-5.6 Sol
34.7%
5
Image
Opus 4.8
30.5%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
GLM 5.3
11.9%
3
Image
Fable 5.1
11.3%
4
Image
Muse Spark 1.3
10.5%
5
Image
Kimi K3
10.2%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Binary Accuracy
1
Image
GPT-6 Astra
34.2%
2
Image
Muse Spark 1.3
32.2%
3
Image
Opus 5
31.6%
4
Image
GPT 5.6 Sol
30.6%
5
Image
GPT 5.6 Luna
30.3%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
4
Image
GLM 5.3
32.4%
5
Image
Grok 4.6
26.5%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Muse Spark 1.3
17.7%
2
Image
Grok 4.6
17.1%
3
Image
Fable 5.1
15.8%
4
Image
Qwen 3.8 Max
14.9%
5
Image
Opus 5
14.9%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Fable 5.1
39.6%
3
Image
Opus 5
38.9%
4
Image
Grok 4.6
33.6%
5
Image
Gemini Flash 3.8
32.6%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By binary accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · xhigh
36.89%
3
Image
Opus 5 · high
33.33%
4
Image
Opus 5 · medium
25.33%
5
Image
Opus 5 · xhigh (older run)
30.21%
Scientific & Research Workflows

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
Fable 5.1
40.0%
2
Image
Opus 5
30.0%
3
Image
GPT-5.6 Sol
22.4%
4
Image
Claude Fable 5
21.4%
5
Image
DeepSeek V4.1 Flash
15.7%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
4
Image
Muse Spark 1.3
12.6%
5
Image
Claude Opus 5
12.2%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Fable 5.1
19.6%
4
Image
Kimi K3
17.3%
5
Image
Claude Opus 5
16.5%
Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Fable 5.1
14.5%
2
Image
Opus 5
14.0%
3
Image
GPT-6 Astra
12.4%
4
Image
Gemini Flash 3.8
6.3%
5
Image
Grok 4.6
5.7%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
GLM 5.3
30.4%
2
Image
DeepSeek V4 Pro
28.5%
3
Image
Kimi K3
27.9%
4
Image
Fable 5.1
24.0%
5
Image
GPT-6 Astra
21.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

Top Models by Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3-Codex
26.02%
3
Image
GPT-5.4
23.47%
4
Image
GPT 5.2
21.94%
5
Image
Opus 4.6
20.92%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
4
Image
Kimi K3
36.1%
5
Image
Fable 5.1
35.9%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

Top Systems (Agg. Reward)
1
Image
Sonnet 4.6 · ICL
+0.196
2
Image
GPT-5.4 · ICL
+0.189
3
Image
Sonnet 4.6 · Claude Code
+0.185
4
Image
Opus 4.7 · Claude Code
+0.183
5
Image
GPT 5.4 · Mem0
+0.148
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
4
Image
Fable 5
44.5%
5
Image
GLM-5.3
41.8%
of
The Frontier AI Data Lab

Data development for the frontier

Snorkel partners with frontier AI teams to build the data, evaluation systems, and environments to improve models where generic coverage runs out.

Snorkel Data Series

Ready-to-use curriculum-structured datasets for the task areas frontier models are pushing hardest, with rubrics, reviewer guidance, difficulty tiers, and eval slices built in.

Custom data development

When off-the-shelf coverage runs out, we build bespoke datasets, evals, and benchmark expansions for the exact failure surface you need to close.

Specialized agents

Card content
Data

Expert Demonstrations & Reasoning

Human solution traces
Reasoning traces
SME Q&A rationales
Workflow demos and decision workflows
Tool-use demos

Preference Labels & Rankings

Patch/draft/report quality ranking
Trajectory QA
Risk/safety/style calibration
Helpful/harmless ranking
Grounding & style

Rubrics & Verifiable Outcomes

Unit tests / compile
Deterministic graders
Citation correctness
Numerical consistency/scorable math/science
Long-horizon tasks
Environments

Standard & Custom Environments

Repo + CLI tools
Browser/GUI harness
Multi-step/stateful workflows
Simulated environments
Your tools, codebase, corpus, data & permissions
DATA DEVELOPMENT

Good data is a set of design choices

Most data quality problems are design problems. Ambiguous task definitions produce inconsistent labels. Uncalibrated reviewers introduce systematic bias. Missing provenance makes failure analysis guesswork. Snorkel's proprietary process is built around the decisions that determine whether training data actually drives model improvement:
Custom AGENTS

Specialized agents grounded in expert data

The same data development system we use to improve frontier models powers our specialized agents. That means agents evaluated against task-specific rubrics and programmatic checks – not generic benchmarks – and refined through the same adjudication and provenance practices used in production model development.

Image
Built for specialized workflows and high-consequence decisions, not generic copilots
Image
Evaluation on environment-grounded tasks with programmatic pass/fail criteria
Image
Same rigor used to train frontier-class models, applied to your enterprise deployment
PUBLISHED RESEARCH

Research that shapes the work

Every dataset, benchmark, and environment we create is the output of active research co-developed and peer-reviewed with leading academic teams and frontier labs.

Image
Image

For models that need to be right. Not just good enough.