workplace agents //
GDPVal+
Train & evaluate frontier agents on the professional work the economy runs on
GDPval+ is Snorkel’s data series for training and evaluating whether AI can do a broad set of professional jobs across domains, roles, and industries.
Developed by Snorkel's AI Data Research Lab, GDPval+ delivers longer-horizon tasks that produces tangible deliverables like a document, spreadsheet or presentation, drawn from real workflows. With domain expert-curated tasks across all 20 O*NET sectors and 100+ occupations, you can cover up to 100% of the US labor market.
Sector coverage includes
- Manufacturing: Plant operations, supply chain, and engineering deliverables
- Professional, Scientific & Technical Services: Legal, market research, microbiology, and information security work products
- Health Care & Social Assistance: Clinical formulations, authorization packages, and care-team workflows
- Educational Services: Curriculum design, assessment, and instructional materials
- Construction: Project planning, site inspection, and compliance deliverables
Other Services: Repair, personal services, and civic-organization workflows
Public Administration: Government and emergency-management deliverables
- Retail Trade: Retail operations, merchandising, and customer-service workflows
- Transportation & Warehousing: Logistics planning, dispatch, and warehouse-operations tasks
Arts, Entertainment & Recreation: Creative production and venue-operations workflows
Plus 10 more sectors covering the rest of the U.S. digital labor market.
GDPval+ is intentionally calibrated to stress-test state-of-the-art agents
Built for frontier model evaluation:
- Tiered difficulty from Core to Frontier
- Frontier tasks where leading models score ≤20% accuracy
- Evaluated against GPT-5.4-Thinking and Claude Opus 4.6 in the Harbor harness, graded by Gemini 2.5 Pro
If your agent succeeds here, it can do the work of an industry professional.
Why the Snorkel Data Series
Expert-led validation
Every task is built and validated through a multi-layer quality pipeline.
Expert review
Expert contributors author every task; subject-matter experts review each one against acceptance criteria and metadata accuracy.
Programmatic checks
Automated validation ensures task uniqueness, minimum resource requirements, and rubric quality.
Difficulty validation
Task difficulty labels are validated against observed accuracy from a panel of frontier models.
Distribution guardrails
New submissions are accepted only if they maintain dataset balance across task types, difficulty levels, and categories.