resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers and more.
Image for Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark
Blog

Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark

Announcing a $3M commitment to launch Open Benchmarks Grants
September 30, 2025
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory
Blog

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

Announcing a $3M commitment to launch Open Benchmarks Grants
March 31, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
of
Type: All Types
Sort: Newest
On the Opportunities and Risks of Foundation Models
Stanford researchers concluded that new, larger and more powerful foundation models represent a paradigm shift in AI, providing opportunities and risks that require deep interdisciplinary collaboration to understand and address.
Research Paper
On the Opportunities and Risks of Foundation Models

Stanford researchers concluded that new, larger and more powerful foundation models represent a paradigm shift in AI, providing opportunities and risks that require deep interdisciplinary collaboration to understand and address.

Mar 15, 2023
Snorkel Team
Learn more about On the Opportunities and Risks of Foundation Models
Ask Me Anything: A simple strategy for prompting language models.
This paper proposes "Ask Me Anything" (AMA), a prompting method that uses weak supervision to combine noisy predictions from multiple prompts generated from an LLM, resulting in an average 10.2% performance lift over the few-shot baseline across a variety of different open-source models.
Research Paper
Ask Me Anything: A simple strategy for prompting language models.

This paper proposes “Ask Me Anything” (AMA), a prompting method that uses weak supervision to combine noisy predictions from multiple prompts generated from an LLM, resulting in an average 10.2% performance lift over the few-shot baseline across a variety of different open-source models.

Mar 15, 2023

S. Arora, et al.

Learn more about Ask Me Anything: A simple strategy for prompting language models.
Contrastive Adapters for Foundation Model Group Robustness
The authors propose Contrastive Adapting, an efficient adapter training strategy that improves the group robustness of large pretrained foundation models (FMs) without finetuning, leading to up to 56.0 percentage points of increase in accuracy compared to zero-shot.
Research Paper
Contrastive Adapters for Foundation Model Group Robustness

The authors propose Contrastive Adapting, an efficient adapter training strategy that improves the group robustness of large pretrained foundation models (FMs) without finetuning, leading to up to 56.0 percentage points of increase in accuracy compared to zero-shot.

Mar 15, 2023

M. Zhang, et al.

Learn more about Contrastive Adapters for Foundation Model Group Robustness
Zero-Shot Learning with Common Sense Knowledge Graphs
Zero-shot learning with Common Sense Knowledge Graphs is a general-purpose framework with a novel transformer graph convolutional network for generating class representations from common sense knowledge graphs, which improves over existing WordNet-based methods on zero-shot learning tasks.
Research Paper
Zero-Shot Learning with Common Sense Knowledge Graphs

Zero-shot learning with Common Sense Knowledge Graphs is a general-purpose framework with a novel transformer graph convolutional network for generating class representations from common sense knowledge graphs, which improves over existing WordNet-based methods on zero-shot learning tasks.

Mar 15, 2023
Snorkel Team
Learn more about Zero-Shot Learning with Common Sense Knowledge Graphs
Binary Classification with Positive Labeling Sources
This paper demonstrates that WEAPO, a Weak Supervision method for binary classification tasks with only positive labeling sources, is effective and efficient—achieving the highest performance of the tested Weak Supervision approaches in terms of label quality and final classifier accuracy on 10 benchmark datasets.
Research Paper
Binary Classification with Positive Labeling Sources

This paper demonstrates that WEAPO, a Weak Supervision method for binary classification tasks with only positive labeling sources, is effective and efficient—achieving the highest performance of the tested Weak Supervision approaches in terms of label quality and final classifier accuracy on 10 benchmark datasets.

Mar 15, 2023

J. Zhang, et al.

Learn more about Binary Classification with Positive Labeling Sources
Tight Lower Bounds on Worst-Case Guarantees for Zero-Shot Learning with Attributes
This paper demonstrates a mathematical analysis of zero-shot learning with attributes, providing a tight lower bound on the worst-case error of the best map from attributes to classes and showing that this bound is predictive of how standard zero-shot methods behave in practice.
Research Paper
Tight Lower Bounds on Worst-Case Guarantees for Zero-Shot Learning with Attributes

This paper demonstrates a mathematical analysis of zero-shot learning with attributes, providing a tight lower bound on the worst-case error of the best map from attributes to classes and showing that this bound is predictive of how standard zero-shot methods behave in practice.

Mar 15, 2023

A. Mazzetto, et al.

Learn more about Tight Lower Bounds on Worst-Case Guarantees for Zero-Shot Learning with Attributes
AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels
AutoWS-Bench-101 is a framework for evaluating automated weak supervision techniques compared to other baseline methods such as zero-shot foundation models and supervised learning, in order to help practitioners choose the best method to generate additional labels.
Research Paper
AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels

AutoWS-Bench-101 is a framework for evaluating automated weak supervision techniques compared to other baseline methods such as zero-shot foundation models and supervised learning, in order to help practitioners choose the best method to generate additional labels.

Mar 15, 2023
Snorkel Team
Learn more about AutoWS-Bench-101: Benchmarking Automated Weak Supervision with 100 Labels
Lifting Weak Supervision To Structured Prediction
This paper finds that weak supervision can be used beyond classification applications, including rankings, graphs, and manifolds, and can provide generalization guarantees nearly identical to models trained on clean data.
Research Paper
Lifting Weak Supervision To Structured Prediction

This paper finds that weak supervision can be used beyond classification applications, including rankings, graphs, and manifolds, and can provide generalization guarantees nearly identical to models trained on clean data.

Mar 15, 2023

Vishwakarma, et al

Learn more about Lifting Weak Supervision To Structured Prediction
Understanding Programmatic Weak Supervision via Source-aware Influence Function
This paper proposes source-aware variation of Influence Function, which measures the influence of individual components in the Programmatic Weak Supervision pipeline, and can be used for multiple purposes such as understanding incorrect predictions, identifying mislabeling of sources, and improving the end model's generalization performance.
Research Paper
Understanding Programmatic Weak Supervision via Source-aware Influence Function

This paper proposes source-aware variation of Influence Function, which measures the influence of individual components in the Programmatic Weak Supervision pipeline, and can be used for multiple purposes such as understanding incorrect predictions, identifying mislabeling of sources, and improving the end model’s generalization performance.

Mar 15, 2023

J. Zhang, et al

Learn more about Understanding Programmatic Weak Supervision via Source-aware Influence Function
1 38 39 40 65
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.