resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers and more.
Image for Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark
Blog

Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark

Announcing a $3M commitment to launch Open Benchmarks Grants
September 30, 2025
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory
Blog

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

Announcing a $3M commitment to launch Open Benchmarks Grants
March 31, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
of
Type: All Types
Sort: Newest
The future of large language models is faster and more robust
Blog
The future of large language models is faster and more robust

Snorkel and affiliated academic labs have been hard at work reducing how computationally expensive large language models are.

Jun 29, 2023
Learn more about The future of large language models is faster and more robust
Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias
Large language models (LLMs) have been recently leveraged as training data generators for various natural language processing (NLP) tasks. While previous research has explored different approaches to training models using generated data, they generally rely on simple class-conditional prompts, which may limit the diversity of the generated data and inherit systematic biases of LLM. Thus, we investigate training data generation with diversely attributed prompts (e.g., specifying attributes like length and style), which have the potential to yield diverse and attributed generated data. Our investigation focuses on datasets with high cardinality and diverse domains, wherein we demonstrate that attributed prompts outperform...
Research Paper
Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias

Large language models (LLMs) have been recently leveraged as training data generators for various natural language processing (NLP) tasks. While previous research has explored different approaches to training models using generated data, they generally rely on simple class-conditional prompts, which may limit the diversity of the generated data and inherit systematic biases of LLM. Thus, we investigate training data generation…

Jun 28, 2023

Y. Yu, et al.

Learn more about Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias
Claypot AI CEO on why you should deploy models the hard way
Blog
Claypot AI CEO on why you should deploy models the hard way

Claypot AI CEO Chip Huyen presented “Platform for Real-Time Machine Learning” at Snorkel AI’s Future of Data-Centric AI 2022.

Jun 27, 2023
Learn more about Claypot AI CEO on why you should deploy models the hard way
LLMs high priority for enterprise data science, but concerns remain
Blog
LLMs high priority for enterprise data science, but concerns remain

Enterprises—especially the world’s largest—are excited to use large language models, but they want to fine-tune them on proprietary data.

Jun 23, 2023
Learn more about LLMs high priority for enterprise data science, but concerns remain
McKinsey QuantumBlack on automating data quality remediation with AI
Blog
McKinsey QuantumBlack on automating data quality remediation with AI

Jacomo Corbo and Bryan Richardson with QuantumBlack present “Automating Data Quality Remediation With AI” at The Future of Data-Centric AI.

Jun 22, 2023
Learn more about McKinsey QuantumBlack on automating data quality remediation with AI
Sambanova on using LLMs to squeeze value from business data
Blog
Sambanova on using LLMs to squeeze value from business data

Stefano Lindt presents “Leveraging NLP to Extract Value From Business Data” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022.

Jun 20, 2023
Learn more about Sambanova on using LLMs to squeeze value from business data
Black Swan Data CTO on how to tackle petabyte-level learning
Blog
Black Swan Data CTO on how to tackle petabyte-level learning

Peter Davio, CTO at Black Swan Data, presented “Petabyte-Level Learning” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022.

Jun 15, 2023
Learn more about Black Swan Data CTO on how to tackle petabyte-level learning
How Grammarly strives for superhuman communication assistance
Blog
How Grammarly strives for superhuman communication assistance

Grammarly’s Timo Mertens presents “Toward Superhuman Communication Assistance” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022.

Jun 14, 2023
Learn more about How Grammarly strives for superhuman communication assistance
The Future of Data-Centric AI Day 2: Snorkel Flow and Beyond
Blog
The Future of Data-Centric AI Day 2: Snorkel Flow and Beyond

The Future of Data-Centric AI showcased customer to successes, took a deep look at Snorkel Flow, and announced two new solutions.

Jun 10, 2023
Learn more about The Future of Data-Centric AI Day 2: Snorkel Flow and Beyond
1 33 34 35 65
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.