Image
author

Zhengyang (Jason) Qi

Research Scientist
,
Snorkel AI

I am an aspiring AI researcher with a diverse range of experience in frontier AI research, large scalable machine learning systems, and applied analytics in social science. I believe in the interactionist approach to intelligence development, through granular feedbacks from grounded, open-ended environments, where robust rewards are essential to forge systems that learn, adapt, and evolve through interactions.

The latest from Jason

Milestone-Based Evaluation and Training for Long-Horizon AI Agents
Blog
Milestone-Based Evaluation and Training for Long-Horizon AI Agents

Long-horizon agents operate across many dependent states and transitions, often spanning multiple tools, environments, and periods of external feedback. The difficulty comes from preserving coherent progress as earlier decisions constrain later actions. A single workflow may involve researching evidence, changing files or records, waiting for external responses, revising plans, validating intermediate results, and returning to earlier systems with new information….

Jul 09, 2026
Learn more about Milestone-Based Evaluation and Training for Long-Horizon AI Agents
Cua-Bench: benchmarking computer-use agents on professional software
Blog
Cua-Bench: benchmarking computer-use agents on professional software

TL;DR We built a benchmark of 25 expert-authored KiCad schematic-editing tasks and ran a frontier computer-use agent against them. The headline numbers: 1. Why build a computer-use benchmark for electrical engineering? Most computer-use benchmarks today live in the same handful of apps: web browsers, file managers, generic productivity suites. Those evaluations are useful, but they share a structural weakness —…

Learn more about Cua-Bench: benchmarking computer-use agents on professional software
Image

For models that need to be right. Not just good enough.