SNORKEL DATA SERIES //
Scientific & Research Workflows //
PaperBench+
Train and evaluate agents on the ability to replicate AI research
PaperBench+ is Snorkel's dataset for evaluating and training agents on the replication and implementation work of an AI researcher.
Developed by Snorkel's AI research team, it builds upon the original PaperBench benchmark and delivers reproduction tasks built from papers published across top-tier venues, each packaged as a self-contained Harbor task with an oracle solution and a methodology-aware rubric.
REQUEST DATA SAMPLES //
Domain coverage
Domains
Language & Reasoning
Vision & Multimodal
Theory, Eval & Interpolation
AI for Science & Math
Graph Learning
Generative Models
Tabular & Classical ML
Contribution types
Architectural innovation
Post-training / alignment
Inference-time techniques
Efficiency / compression
Data engineering
Evaluation / mechanisms / other
Papers are sourced from top-tier venues such as ICML, NeurIPS, ICLR, ACL, EMNLP, and CVPR.
Built for frontier model evaluation and training:
- Tiered difficulty from Core to Frontier, calibrated by frontier-model pass rate
- Frontier tier where models pass fewer than 20% of attempts
- Grading combines LLM judging with deterministic checks
- Every task ships with an oracle solution verified to reproduce the paper's key metrics within reported tolerance