resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers and more.
Image for Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark
Blog

Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark

Announcing a $3M commitment to launch Open Benchmarks Grants
September 30, 2025
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory
Blog

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

Announcing a $3M commitment to launch Open Benchmarks Grants
March 31, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
of
Type: All Types
Sort: Newest
Alfred: Data labeling with foundation models and weak supervision
Blog
Alfred: Data labeling with foundation models and weak supervision

Introducing Alfred: an open-source tool for combining foundation models with weak supervision for faster development of academic data sets.

Aug 27, 2024
Learn more about Alfred: Data labeling with foundation models and weak supervision
Language Models in the Loop: Incorporating Prompting into Weak Supervision
We propose a new strategy for applying large pre-trained language models to novel tasks when labeled training data is limited. Rather than apply the model in a typical zero-shot or few-shot fashion, we treat the model as the basis for labeling functions in a weak supervision framework. To create a classifier, we first prompt the model to answer multiple distinct queries about an example and define how the possible responses should be mapped to votes for labels and abstentions. We then denoise these noisy label sources using the Snorkel system and train an end classifier with the resulting training data....
Research Paper
Language Models in the Loop: Incorporating Prompting into Weak Supervision

We propose a new strategy for applying large pre-trained language models to novel tasks when labeled training data is limited. Rather than apply the model in a typical zero-shot or few-shot fashion, we treat the model as the basis for labeling functions in a weak supervision framework. To create a classifier, we first prompt the model to answer multiple distinct…

Aug 22, 2024

R. Smith et al.

Learn more about Language Models in the Loop: Incorporating Prompting into Weak Supervision
Webinar
Curate training data via labeling functions— 10 to 100x faster

In this webinar, we’ll explain how enterprises can not only accelerate data labeling but iterate, adapt, and improve label accuracy via AI data development.

Aug 22, 2024
Snorkel Team
Learn more about Curate training data via labeling functions— 10 to 100x faster
Webinar
Distilling LLMs into SLMs for higher accuracy and lower inference costs

In this webinar, we’ll provide an overview of LLM distillation, explain how it compares with fine-tuning, and introduce the latest techniques for training SLMs using larger models and knowledge transfer.

Aug 21, 2024
Snorkel Team
Learn more about Distilling LLMs into SLMs for higher accuracy and lower inference costs
RAG: LLM performance boost with retrieval-augmented generation
Blog
RAG: LLM performance boost with retrieval-augmented generation

Retrieval-augmented generation (RAG) enables LLMs to produce more accurate responses by finding and injecting relevant context. Learn how.

Aug 15, 2024
Learn more about RAG: LLM performance boost with retrieval-augmented generation
Call center AI for customer experience management: a case study
Blog
Call center AI for customer experience management: a case study

How one large financial institution used call center AI to inform customer experience management with real-time data.

Aug 14, 2024
Learn more about Call center AI for customer experience management: a case study
An introduction to Snorkel and data-centric AI
eBook
An introduction to Snorkel and data-centric AI

Learn how Snorkel can programmatically help you create massive amounts of high-quality labeled training data in a matter of hours.

Aug 09, 2024
Snorkel Team
Learn more about An introduction to Snorkel and data-centric AI
New GenAI features, data annotation: Snorkel Flow 2024.R2
Blog
New GenAI features, data annotation: Snorkel Flow 2024.R2

This release features new GenAI tools and Multi-Schema Annotation, as well as new enterprise security tools and an updated home page.

Aug 07, 2024
Learn more about New GenAI features, data annotation: Snorkel Flow 2024.R2
Webinar
How to optimize RAG pipelines for domain- and enterprise-specific tasks

RAG is the first step in building LLM-powered AI applications for enterprise use cases.

Aug 01, 2024
Snorkel Team
Learn more about How to optimize RAG pipelines for domain- and enterprise-specific tasks
1 14 15 16 65
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.