ImageImage
ImageImage

SNORKEL DATA SERIES //

Scientific & Research Workflows // 

PaperBench+

Train and evaluate agents on the ability to replicate AI research 

PaperBench+ is Snorkel's dataset for evaluating and training agents on the replication and implementation work of an AI researcher.

Developed by Snorkel's AI research team, it builds upon the original PaperBench benchmark and delivers reproduction tasks built from papers published across top-tier venues, each packaged as a self-contained Harbor task with an oracle solution and a methodology-aware rubric.

REQUEST DATA SAMPLES //
By submitting this form, I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.

Domain coverage

Domains
Language & Reasoning
Vision & Multimodal
Theory, Eval & Interpolation
AI for Science & Math
Graph Learning
Generative Models
Tabular & Classical ML
Contribution types
Architectural innovation
Post-training / alignment
Inference-time techniques
Efficiency / compression
Data engineering
Evaluation / mechanisms / other

Papers are sourced from top-tier venues such as ICML, NeurIPS, ICLR, ACL, EMNLP, and CVPR.

Built for frontier model evaluation and training:

  • Tiered difficulty from Core to Frontier, calibrated by frontier-model pass rate
  • Frontier tier where models pass fewer than 20% of attempts
  • Grading combines LLM judging with deterministic checks
  • Every task ships with an oracle solution verified to reproduce the paper's key metrics within reported tolerance
Image
Image

Train models on the building blocks of scientific discovery with the Snorkel Data Series