SNORKEL DATA SERIES
Agentic Coding
The Agentic Coding Data Series captures every dimension of real software engineering—from multi-step problem solving and iterative debugging to environment manipulation and tool use within a command-line interface.
Developed by Snorkel’s AI Data Research Lab in collaboration with leading experts in software engineering, this extension of Terminal-Bench 2.0 provides high-signal datasets and deterministic evaluation environments that help you build, test, and tune AI systems that perform like real coding experts.
Expert-led validation
Human review — SMEs validate clarity, correctness, and that each task is fully solvable.
LLMaJ validation — Automated checks detect instruction–test mismatches and missing details.
Deterministic testing — Code-based checks to validate benchmark compliance, formatting, syntax, etc.
Guardrails — Additional checks to detect and limit cheating paths, non-deterministic behavior, and reward hacking.
Coding tasks including:
System / environment setup & configuration
Build / Compilation / dependency management
Data / file processing / ETL / scripting
Interactive / simulation tasks / games
Machine learning / model training / inference
Debugging / repair / fixing tasks
Security / cryptography / vulnerability demonstration
Scientific-computing
Why the Snorkel Data Series
High-volume quarterly drops
Multi-layer quality pipeline
Unified execution environment
Direct roadmap influence
Let’s talk
By submitting this form, I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.