Image
Research-led data development

Datasets and environments that give frontier models domain expertise

Snorkel builds the human expert-authored datasets, evaluation environments, and benchmarks calibrated to push the limits of frontier model capability. Off-the-shelf or custom.

Where generic data runs out

Frontier model development stalls on data problems generic pipelines weren't built to solve, including distributional gaps in specialized domains, benchmark blind spots, and failure modes that only surface at scale. We builds the data to solve them.

get started

Two ways to get the data you need

What the data frontier models need most is rarely the data that already exists. Snorkel delivers it two ways: off-the-shelf for well-defined task areas, or custom-built for the gaps only you can see.

Image
Snorkel Data Series
Ready-to-use datasets for the task areas where frontier models are actually being pushed. Each ships with rubrics, reviewer guidance, difficulty tiers, and eval slices built in.
Image
Custom data development

Bespoke datasets, evaluation environments, and benchmark expansions to target the exact failure surface you're trying to close.

SNORKEL DATA SERIES

Datasets and environments

Readily available datasets developed in close collaboration with leading frontier AI teams – curriculum-structured to build difficulty progressively across a task area, with the evaluation infrastructure to match.
Alignment for better code generation icon
Terminal Coding
Long-horizon agents in real containerized terminals — planning, execution, and error recovery.
Terminal-Bench 2.0 & 3.0
Image
Software Engineering
Debugging, codebase understanding, and multi-file changes across 7+ languages.

SWE-Bench Pro, Senior SWE-bench

enterprise environments
Enterprise & Workplace Environments
Multi-turn, tool-rich professional workflows across industries, policies, and 100+ occupations.
τ²-bench, τ³-bench, GDPval
Image
Computer Use
GUI interaction and desktop workflow execution across real applications.
OSWorld 2.0, Agent’s Last Exam
Image
Scientific & Research Workflows
Research execution and technical reasoning across scientific domains.
Terminal-Bench-Science, PaperBench
The frontiers of multi-turn math reasoning  icon
STEM Knowledge & Reasoning
Expert scientific problems requiring multi-step reasoning.
Humanity’s Last Exam, FrontierMath

Featured data series

Terminal coding

Senior SWE-bench+

Real-world PR tasks at a senior engineer’s bar
— inferring intent, runtime debugging, and shipping code that fits the codebase. 

Led benchmark development

Enterprise agents

GDPval+

Economically-grounded professional
tasks spanning all 20 O*NET sectors
and 100+ occupations.

Terminal coding

Frontier-Bench+

Multi-step terminal engineering tasks, calibrated to the Frontier-Bench standard.

Core benchmark contributor and reviewer

icon bulb
For the full Snorkel data catalog, talk to our team.
Image
CUSTOM DATA DEVELOPMENT

Closing gaps existing datasets can’t reach

Custom data development engagements start with the failure surface: what the model can't do, where it's brittle, and what the correct evaluation criteria are. From there, Snorkel builds the datasets, environments, and benchmark expansions needed to close it.

01
Task specification and rubric design
02
Bespoke dataset construction
03
RL environment development
04
Benchmark and eval expansion
05
Provenance and adjudication
Image
Image

For models that need to be right. Not just good enough.