We define and advance data and environments to push the AI frontier

Built on 10+ years of pioneering research in data-centric AI,
including 250+ publications and benchmarks.

building benchmarks and collaborating with

Image
Image
Image
Image
Image
Image
Image
Image
Image
key research areas

Vision and impact

We help labs advance frontier models by working with domain experts to design and build complex, realistic datasets that drive model performance.

initiatives

Community and open science

Open benchmarks, conversations, and research for real-world AI performance.

Image

Open Benchmarks Grants

Backed by a $3M commitment, the program funds
open-source datasets, benchmarks, and evaluation artifacts that shape how frontier AI systems are built
and evaluated.

Image

Bench Talks

Our podcast series at the intersection of AI evaluation, data quality, and real-world impact.
Image

Reading Group

A recurring forum for researchers and practitioners to explore the latest frontier developments in AI while building meaningful connections within the community.

DEEP RESEARCH Expertise

Technical advisors and distinguished affiliates

Stephen Bach headshot

Stephen Bach

Brown University
Eliot Horowitz Assistant Professor, Computer Science Department
Jason Fries headshot

Jason Fries

Stanford University
Assistant Professor of Biomedical Data Science and of Medicine
Jared Dunnmon headshot

Jared Dunnmon

Co-Founder & Chief Scientist, Stealth Startup
Prev. Dir. of AI at DIU
Fred Sala headshot

Fred Sala

Chief Scientist
,
Snorkel AI
Assistant Professor @ University of Wisconsin-Madison
Chris Ré headshot

Chris Ré

Co-Founder
,
Snorkel AI
Professor @ Stanford University
Ludwig Schmidt headshot

Ludwig Schmidt

Stanford University · LAION
Stanford researcher and LAION collaborator
Karthik Narasimhan headshot

Karthik Narasimhan

Princeton University
Professor of Computer Science
Yu Su headshot

Yu Su

Ohio State University
Associate Professor of Computer Science and Engineering
Lewis Tunstall headshot

Lewis Tunstall

Hugging Face
Machine Learning Engineer
PUBLICATIONS

Browse research blogs
and academic papers

Type: All Types
Sort: Newest
Large language models: their history, capabilities and limitations
Blog
Large language models: their history, capabilities and limitations

Large language models have enormous potential. But what are they? Where did they come from? And how can you make them work better?

May 25, 2023
Learn more about Large language models: their history, capabilities and limitations
Stanford professor on data-centric AI for healthcare and medicine
Blog
Stanford professor on data-centric AI for healthcare and medicine

Stanford assistant professor James Zou, presents “Responsible Data-Centric AI for Healthcare and Medicine” at The Future of Data-Centric AI.

May 18, 2023
Learn more about Stanford professor on data-centric AI for healthcare and medicine
Poster presenters compete to win desktop GPU
Blog
Poster presenters compete to win desktop GPU

Snorkel AI has accepted the first batch of applications for its first annual virtual poster competition. But there’s still time to add yours to the mix.

May 09, 2023
Learn more about Poster presenters compete to win desktop GPU
Use your data to build your AI moat: The Future of Data-Centric AI 2023
Blog
Use your data to build your AI moat: The Future of Data-Centric AI 2023

Join us on June 7-8 to learn how to use your data to build your AI moat at The Future of Data-Centric AI 2023 free virtual conference.

May 04, 2023
Learn more about Use your data to build your AI moat: The Future of Data-Centric AI 2023
Out of distribution blindness: why to fix it and how energy can help
Blog
Out of distribution blindness: why to fix it and how energy can help

Sharon Li is an assistant professor at the University of Wisconsin-Madison. She presented “Detecting Data Distributional Shift: Challenges and Opportunities” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022. The talk covered a novel approach for handling out-of-distribution objects.

May 03, 2023
Learn more about Out of distribution blindness: why to fix it and how energy can help
Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine
Research Paper
Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine
May 01, 2023

J. Lemmon, et al.

Learn more about Evaluation of Feature Selection Methods for Preserving Machine Learning Performance in the Presence of Temporal Dataset Shift in Clinical Medicine
Harvard professor: DataPerf and AI’s need for data benchmarks
Blog
Harvard professor: DataPerf and AI’s need for data benchmarks

Harvard Professor Vijay Janapa Reddi’s presentation: “DataPerf: Benchmarks for data” from Snorkel AI’s 2022 Future of Data-Centric AI event.

Apr 25, 2023
Learn more about Harvard professor: DataPerf and AI’s need for data benchmarks
Learning to Compose Soft Prompts for Compositional Zero-Shot Learning
We introduce compositional soft prompting (CSP), a parameter-efficient learning technique to improve the zero-shot compositionality of large-scale pretrained vision-language models (VLMs) like CLIP. We develop CSP for compositional zero-shot learning, the task of predicting unseen attribute-object compositions (e.g., old cat and young tiger). VLMs have a flexible text encoder that can represent arbitrary classes as natural language prompts but they often underperform taskspecific architectures on the compositional zero-shot benchmark datasets. CSP treats the attributes and objects that define classes as learnable tokens of vocabulary. During training, the vocabulary is tuned to recognize classes that compose tokens in multiple ways (e.g.,...
Research Paper
Learning to Compose Soft Prompts for Compositional Zero-Shot Learning

We introduce compositional soft prompting (CSP), a parameter-efficient learning technique to improve the zero-shot compositionality of large-scale pretrained vision-language models (VLMs) like CLIP. We develop CSP for compositional zero-shot learning, the task of predicting unseen attribute-object compositions (e.g., old cat and young tiger). VLMs have a flexible text encoder that can represent arbitrary classes as natural language prompts but they…

Apr 24, 2023

N. Nayak et al.

Learn more about Learning to Compose Soft Prompts for Compositional Zero-Shot Learning
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
A long standing goal of the data management community is to develop general, automated systems that ingest semi-structured documents and output queryable tables without human effort or domain specific customization. Given the sheer variety of potential documents, state-of-the art systems make simplifying assumptions and use domain specific training. In this work, we ask whether we can maintain generality by using large language models (LLMs). LLMs, which are pretrained on broad data, can perform diverse downstream tasks simply conditioned on natural language task descriptions. We propose and evaluate EVAPORATE, a simple, prototype system powered by LLMs. We identify two fundamentally different...
Research Paper
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes

A long standing goal of the data management community is to develop general, automated systems that ingest semi-structured documents and output queryable tables without human effort or domain specific customization. Given the sheer variety of potential documents, state-of-the art systems make simplifying assumptions and use domain specific training. In this work, we ask whether we can maintain generality by using…

Apr 21, 2023

S. Arora, et al.

Learn more about Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
1 16 17 18 35
Image

Let’s research together

Join our team of leading researchers and help shape the future of AI.