ImageImage
ImageImage

SNORKEL DATA SERIES //

Specialized Computer Use Agents //

CUA-Bench+

Stress-test & train frontier agents on real-world GUI workflows

CUA-Bench provides a unified, cross-platform benchmark for evaluating and training agents on realistic GUI-based work, including mouse clicks, keyboard inputs, and screen reading within dynamic desktop and mobile environments.

Developed by Snorkel's AI Data Research Lab, CUA-Bench+ delivers thousands of verifiable tasks across Linux, Windows, macOS, and Android — built to test multi-step reasoning, error recovery, and multi-application workflows.

REQUEST DATA SAMPLES //
By submitting this form, I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.

Task categories include

  • Artifact Creation
    Creating a new structured professional artifact from scratch

  • Artifact Modification
    Editing an existing artifact to meet specific structural or content constraints

  • Structured Export & Conversion
    Exporting or converting artifacts into a specified deterministic format

  • Project Assembly
    Constructing multi-file project structures with correct relationships and references

  • Cross-Tool Integration
    Producing consistent artifacts across multiple applications in a coordinated workflow

  • Domain-Specific Structured Output
    Producing structured domain-specific definitions that are internally validated by file struture

 

Domains include Electrical Engineering, Computer Science, Architecture, Semiconductors, Bioinformatics, Robotics, Laboratory Information Systems (LIMS), Game Development, Enterprise ERP Systems, 3D Modeling and Animation, Energy and Power Systems, GIS / Geospatial Systems, Manufacturing and Design and more.

CUA-Bench+ is intentionally calibrated to stress-test state-of-the-art agents

Built for frontier model evaluation:

  • Tiered difficulty across Basic, Core, Advanced, and Complex
  • Complex tasks where Claude Opus 4.6 passes less than 20% of attempts
  • Pass rates measured over 5 attempts, with artifact-based verification

Success is scored from the saved output artifact, not the action trace.

Why the Snorkel Data Series

Image
High-volume quarterly drops
Image
Multi-layer quality pipeline
Image
Unified execution environment
Image
Direct roadmap influence

Expert-led validation

Every task is built and validated through a multi-layer quality pipeline.

01

Expert review

Expert contributors author every task; subject-matter experts verify accuracy, clarity, and environment integrity.

02

Programmatic validation

Automated checks ensure manifest completeness, deterministic pass rates, and task uniqueness.

03

Difficulty calibration

Difficulty labels are validated against performance data from frontier LLMs.

04

Distribution guardrails

Submissions are filtered to maintain balanced distribution across categories, platforms, and complexity levels.

Image
Image

Train agents that can navigate GUIs and produce real professional artifacts with the Snorkel Data Series