Research

Data 2.0 and the research era of AI data

September 20, 2026
12 min read

Table of contents

    Introduction

    Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, OnePrime, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, BlackRock, Walden Catalyst Ventures, and Factory.

    Over the last year since launching our new Data-as-a-service offering, we’ve grown over 18x, crossing an annualized revenue run rate of $375M, and we’re honored to now partner with the leading frontier labs, hyperscalars, neolabs, vertical AI leaders, enterprises, and US government agencies who view data as one of the most important ingredients for safe and effective AI.

    We started Snorkel as a research project a decade ago at Stanford. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development merited proper study as a true research and technology problem, not just a staffing and crowdsourcing one.

    Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come.

    At Snorkel, we are building the frontier data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier.

    We are also doubling down on our commitments to support open data development for benchmarking and evaluation, to help guide the increasingly critical path of AI data development; and to support an increasingly diverse ecosystem of general and specialized intelligence.

    Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, and open way. We are excited to support this mission in the next decade of research ahead at Snorkel.

    More detailed thoughts below:

    Data 2.0 – the research lab era of AI data

    All modern automation approaches – digital LLMs, self-driving, physical AI – follow an asymptotic improvement curve where the first mile is driven by volume of simpler data, and the last mile is driven by quality of more complex data. Intuitively: when a model knows very little, almost any data contains additive bits of information, and volume is the key. When a model is expert level, only the right data at the right level of complexity will move the needle.

    The first mile phase – “Data 1.0” – is all about volume, and is largely a staffing and logistics problem. 

    The enduring last mile phase – “Data 2.0” – is all about the right curriculum of extremely complex, high quality data – and is mostly a research and technology problem.

    As a concrete example, consider coding data. Several years ago, LLMs could barely do code auto-complete, and the focus was on large volumes of raw coding data for pre-training, and then, large volumes of simple coding problems and code preference labels for scaled SFT and RLHF.

    Today: LLMs are approaching superhuman capabilities in many areas of coding. Frontier data is more valuable than ever given the economic impacts – but harder than ever to produce. Good coding datasets and environments for RL must approximate software development problems that a senior engineer might struggle with over days or weeks; must distributionally target nuanced model error modes, like a finely calibrated curriculum for an advanced student; must be robust to these advanced models’ attempts to hack, cheat, or circumvent; and often must pass several hundred other quality control checks to be viable.

    The same pattern holds across an exploding surface area of domains AI is tackling, from legal to finance to biology and beyond. Where once getting a sufficient volume of relevant experts to develop basic chatbot tasks or preference labels was enough, now frontier data points and environments represent advanced expert tasks that might take humans days or weeks to complete, have nuanced failure modes or routes for AI cheating, require complex environments simulating e.g. entire companies, and are beyond the ability of even the smartest human experts to develop alone at scale.

    Frontier data is no longer solvable by the optimized supply of human hours and bodies. The frontier ahead will be driven by more elegant research-driven approaches combining the best of human expertise and AI acceleration.

    Data as a human and machine endeavor

    While humans alone will increasingly struggle to develop frontier data, humans must remain at the center of the AI data development loop – accelerated and improved by AI.

    Frontier data must be human-centric, first of all, because frontier data is valuable only insomuch as it targets something that a model does not already know. While technically possible for an LLM to generate good data of this type, purely synthetic data will naturally be highly correlated with what AI models already know – not what they still need to learn. The basic circularity and mode collapse of a purely synthetic approach (other than for distillation) mandates that human-in-the-loop data will be the most valuable kind.

    Second, from a normative standpoint: data is how we measure AI and align it to human values and judgement. Therefore, humans must be involved in the development of data used for AI measurement, alignment, and safety.

    However: the new frontier of data has become far too complex for humans alone to develop. Modern RL environments and datasets now simulate tasks that might take a human days or weeks to accomplish; require hundreds of quality control checks to avoid subtle errors leading to misaligned AI; and must stump the latest advanced models.

    The key challenge, then, is to build AI systems that accelerate and improve human expertise – and in doing so, enable humans to stay at the center of AI data evaluation and development as model capabilities exponentially advance in the decades ahead.

    [Stylized graphic: Human + AI collaboration]

    This is a deep technical project with years of rich work ahead covering many distinct angles. We’ve studied this problem academically for a decade, and built our internal Agentic Data Platform at Snorkel to support it, but there are years of exciting research and development problems ahead, studying how specialized AI models and agents can

    • Accelerate human effort by expanding seeds, constraints, and/or sketches from experts into properly constructed environments and data instances.
    • Guide human efforts towards distributional targets and model error modes.
    • Give live feedback and to do post-submission quality control, review, and revision, in collaboration with human experts.
    • Route the right environment or datapoint subcomponents to the right human experts for review.

    And so on across a rich and growing surface area of human-computer interaction.  

    Pushing beyond trivial synthetic generation to optimally preserve, accelerate, and improve human expertise with AI is one of the most interesting and exciting directions of AI research today (in our slightly biased opinion). The end goal being data that is actually additive to the frontier of AI – and a process of building it that keeps humans in the driver’s seat even as that frontier rapidly accelerates.

    Building the RSI engine for data development

    The most exciting consequence of a human-AI system is its ability to drive a recursive self improvement loop for data, with specialized AI models improving human outcomes, and human supervision in turn improving these models – creating a powerful compounding loop to keep pace with an accelerating RSI model frontier.

    We view building this RSI data engine as one of our most critical objectives. Data is what we use not just to advance AI capabilities, but to measure, monitor, and safely align them – so human-centric data development must keep pace with accelerating RSI model development in this way.

    Much of our work over the last decade of research and development has focused on this critical loop, and our Agentic Data Platform is designed primarily to harvest this core dynamic. Human experts are supported by a stable of specialized AI models and agents – sometimes hundreds per task type – which accelerate and improve human outcomes to keep pace with the advancing frontier, as described in the previous session. In turn, human feedback at scale is used as weak supervision to continually measure and improve these agents. 

    As one example: when building coding agent environments and data, we use hundreds of specialized agents to do quality control in addition to human expert review, which accelerates our QC efficiency by 50%+ and improves accuracy of review by 15+ accuracy points compared to a human review-only baseline. Scaled human review and feedback is then used as weak supervision to improve these agents, resulting in a 2x+ accuracy improvement over a non-specialized frontier LLM baseline. AI improves human experts, and human experts improve AI – and statistics like these continuously improve as the flywheel continues.

    The RSI loop for data development is what has powered our growth to date – allowing exceptionally high data quality in exceptionally complex areas. This will compound and accelerate in the years ahead, allowing human-in-the-loop data to keep pace with the RSI-fueled acceleration of frontier AI.

    Guiding the way with open benchmarks

    We believe that as data and AI development both accelerate, it is more important than ever that this development is guided and measured by an ecosystem of open, independent, and robust benchmarks – all driven by data and environment development.

    Benchmarks are tests for AI, that are essentially collections of datasets and environments that measure progress in key capability areas, and serve as both leaderboards and guideposts for AI progress. In the history of AI, benchmarks have always played an outsized role in motivating and guiding progress. And at their best, they not only motivate and guide development as targets, but provide transparency and insight into model performance, promote safer usage and deployment, and form the backbone of empirical computer science.

    While benchmarks sometimes get pushback as becoming targets for gamification and overfitting, we believe that the best mitigation for this “benchmaxxing” failure mode is an ecosystem of more benchmarks, both public and private, that are robust, diverse, continuous, and independently created. Goodhart’s law – the famous adage that any metric which becomes a target ceases to be a good metric – did not imply that we should drop all metrics, but rather make them more diverse, robust, and less gameable; similarly with benchmarks, the way forward is more, not less.

    At Snorkel, we believe that a good majority of public benchmarks should be developed in open, independent ways, in order to support neutral and diverse evaluation of AI capabilities and gaps. To forward this, we will be doubling down on our Open Benchmarks Grants program, which provides funding, research, and data development support to open, independent benchmarks, and has supported projects like Terminal Bench, OSWorld 2.0, ALE, and many more. More news here very soon!

    Supporting a diverse ecosystem of intelligence

    We believe that the AI ecosystem will ultimately evolve to become a rich one, containing both massive generalist models at the frontier, and a myriad of specialized models tuned for every enterprise, organization, and perhaps even every individual.

    At each level of generalist to specialized, different data and environments are needed. Where generalist frontier models win via maximal coverage of an area, specialized models win via deep focus on specific workflows, use cases, environments, and existing user data or signals that serve as starting points and anchors for data development.

    We believe that an increasingly large portion of the world’s data development will focus on supporting specialized AI over the next several years, leading to an increasingly diverse ecosystem of intelligence – and plan to invest heavily in support of this at Snorkel.

    Supporting safe intelligence

    AI safety is one of the most critical considerations in the days and years ahead – and we believe that safe intelligence fundamentally begins with robust, high quality data.

    Datasets and environments not only drive how we benchmark and evaluate models for safety; they define in a very literal way the objectives and penalties that shape AI model learning and alignment during training. Strong alignment comes from high-quality, well-designed data and environments that are realistic; comprehensive and diverse in their coverage of real world scenarios; and robust to increasingly advanced AI capabilities for cheating or “reward hacking”.  Just as a well-aligned human starts with what they are taught at home, so too does a well aligned, safe AI model start with the data it is trained on.

    A significant portion of our research and development efforts in the years ahead will be focused on developing data and environments that are specifically built to support robust alignment and safe AI.

    The new frontier of AI research

    As AI capabilities rapidly improve, we believe that advancing the science and technology of data and environment development will be one of the most interesting and impactful areas of AI research. The challenges ahead – defining the shape of new Data 2.0 workloads; accelerating the critical collaboration of human experts and AI, and shaping this interaction into an RSI data engine; defining the open benchmarks that guide the field; diversifying to support an expanding ecosystem of intelligence; and supporting fundamentally safe and aligned intelligence – will be some of the most central to AI progress.

    We are incredibly excited for the next decade of research ahead at Snorkel!

    Share this article
    Image
    Alex Ratner
    Co-Founder & CEO

    Alex Ratner is the co-founder and CEO at Snorkel AI, and an affiliate assistant professor of computer science at the University of Washington. Prior to Snorkel AI and UW, he completed his Ph.D. in computer science advised by Christopher Ré at Stanford, where he started and led the Snorkel open source project. His research focused on data-centric AI, applying data management and statistical learning techniques to AI data development and curation.

    Recommended articles

    View all articles
    Image
    Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
    Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8. Benchmark Senior SWE-Bench evaluates agents as compared to a senior engineer.
    September 21, 2026
    Snorkel Team
    Image
    From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
    A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later,
    September 15, 2026
    Snorkel Team
    Image
    Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
    The speed of new frontier model releases keeps accelerating. Meanwhile benchmarks struggle to keep up and saturate quickly, often being left in the dust. Most benchmarks are static datasets with no active maintenance, causing them to lose value fast. Some benchmarks are looking to change this by becoming Continuous Benchmarks. Terminal-Bench is one of the most widely reported benchmarks on
    August 27, 2026
    Justin Bauer
    Image
    Image

    Join our newsletter

    For expert advice, the latest research, and exclusive events.
    By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.