SNORKEL DATA SERIES //
Specialized Computer Use Agents //
CUA-Bench+
Stress-test & train frontier agents on real-world GUI workflows
CUA-Bench provides a unified, cross-platform benchmark for evaluating and training agents on realistic GUI-based work, including mouse clicks, keyboard inputs, and screen reading within dynamic desktop and mobile environments.
Developed by Snorkel's AI Data Research Lab, CUA-Bench+ delivers thousands of verifiable tasks across Linux, Windows, macOS, and Android — built to test multi-step reasoning, error recovery, and multi-application workflows.
Task categories include
Artifact Creation
Creating a new structured professional artifact from scratchArtifact Modification
Editing an existing artifact to meet specific structural or content constraintsStructured Export & Conversion
Exporting or converting artifacts into a specified deterministic formatProject Assembly
Constructing multi-file project structures with correct relationships and referencesCross-Tool Integration
Producing consistent artifacts across multiple applications in a coordinated workflowDomain-Specific Structured Output
Producing structured domain-specific definitions that are internally validated by file struture
Domains include Electrical Engineering, Computer Science, Architecture, Semiconductors, Bioinformatics, Robotics, Laboratory Information Systems (LIMS), Game Development, Enterprise ERP Systems, 3D Modeling and Animation, Energy and Power Systems, GIS / Geospatial Systems, Manufacturing and Design and more.
CUA-Bench+ is intentionally calibrated to stress-test state-of-the-art agents
Built for frontier model evaluation:
- Tiered difficulty across Basic, Core, Advanced, and Complex
- Complex tasks where Claude Opus 4.6 passes less than 20% of attempts
- Pass rates measured over 5 attempts, with artifact-based verification
Success is scored from the saved output artifact, not the action trace.
Why the Snorkel Data Series
Expert-led validation
Every task is built and validated through a multi-layer quality pipeline.
Expert review
Expert contributors author every task; subject-matter experts verify accuracy, clarity, and environment integrity.
Programmatic validation
Automated checks ensure manifest completeness, deterministic pass rates, and task uniqueness.
Difficulty calibration
Difficulty labels are validated against performance data from frontier LLMs.
Distribution guardrails
Submissions are filtered to maintain balanced distribution across categories, platforms, and complexity levels.