resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers and more.
Image for Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark
Blog

Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark

Announcing a $3M commitment to launch Open Benchmarks Grants
September 30, 2025
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory
Blog

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

Announcing a $3M commitment to launch Open Benchmarks Grants
March 31, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
of
Type: All Types
Sort: Newest
Weak Supervision Enables Scalable Post-Market Surveillance on Medical Wearables
Introduction: With the advent of consumer-facing devices that can render atrial fibrillation (AF) pre-diagnosis, medical wearables now have the potential to affect diagnosis rates and medical care. Post-market surveillance is necessary to understand the impact of wearables on patient outcomes and health care utilization, but is hindered by the lack of codified terms in EHR that capture wearable use. Research Questions: Constructing a post-market surveillance system therefore requires a classifier that identifies mentions of AF pre-diagnosis in unstructured EHR data. However, fine-tuning classifiers require large, hand-labeled training sets that can be costly to generate. It is unclear whether a scalable...
Research Paper
Weak Supervision Enables Scalable Post-Market Surveillance on Medical Wearables

Introduction: With the advent of consumer-facing devices that can render atrial fibrillation (AF) pre-diagnosis, medical wearables now have the potential to affect diagnosis rates and medical care. Post-market surveillance is necessary to understand the impact of wearables on patient outcomes and health care utilization, but is hindered by the lack of codified terms in EHR that capture wearable use. Research…

Nov 06, 2023

RM. Yoo, et al.

Learn more about Weak Supervision Enables Scalable Post-Market Surveillance on Medical Wearables
How to fine-tune large language models for enterprise use cases
Blog
How to fine-tune large language models for enterprise use cases

LLMs have a broad but shallow knowledge, but fall short on specialized tasks. For best performance, enterprises must fine tune their LLMs.

Nov 02, 2023
Learn more about How to fine-tune large language models for enterprise use cases
Snorkel Flow 2023.R3 release: PaLM integration, streamlined onboarding, and enhanced user experience
Blog
Snorkel Flow 2023.R3 release: PaLM integration, streamlined onboarding, and enhanced user experience

The 2023.R3 Snorkel Flow release is packed with improvements that amplify user experience, streamline workflows, and enhance performance, ensuring our users derive unparalleled value from our platform.

Nov 01, 2023
Learn more about Snorkel Flow 2023.R3 release: PaLM integration, streamlined onboarding, and enhanced user experience
Navigating Biden’s AI executive order with AI data development
Blog
Navigating Biden’s AI executive order with AI data development

The Biden administration issued an executive order that creates new AI standards and challenges. AI data development can help.

Oct 31, 2023
Learn more about Navigating Biden’s AI executive order with AI data development
Snorkel AI researchers present 18 papers at NeurIPS 2023
Blog
Snorkel AI researchers present 18 papers at NeurIPS 2023

The Snorkel AI team will present 18 research papers and talks at the 2023 Neural Information Processing Systems (NeurIPS) conference from December 10-16. The Snorkel papers cover a broad range of topics including fairness, semi-supervised learning, large language models (LLMs), and domain-specific models. Snorkel AI is proud of its roots in the research community and endeavors to remain at the forefront…

Oct 31, 2023
Learn more about Snorkel AI researchers present 18 papers at NeurIPS 2023
Two approaches to distill LLMs for better enterprise value
Blog
Two approaches to distill LLMs for better enterprise value

Distillation techniques allow enterprises to access the full predictive power of large language models at a tiny fraction of their cost.

Oct 31, 2023
Learn more about Two approaches to distill LLMs for better enterprise value
Enterprise LLM Summit highlights the importance of data development
Blog
Enterprise LLM Summit highlights the importance of data development

Snorkel AI’s Enterprise LLM Virtual Summit drew 1,000 attendees with speakers from Contextual AI, Google, Meta, Stanford, and Together AI.

Oct 27, 2023
Learn more about Enterprise LLM Summit highlights the importance of data development
Zero-Shot Robustification of Zero-Shot Models with Foundation Models
Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their performance. The traditional solution is fine-tuning, but this undermines the key advantage of pretrained models, which is their ability to be used out-of-the-box. We propose ROBOSHOT, a method that improves the robustness of pretrained model embeddings in a fully zero-shot fashion. First, we use zero-shot language models (LMs) to obtain useful insights from task descriptions. These insights are embedded and used to remove harmful and boost useful...
Research Paper
Zero-Shot Robustification of Zero-Shot Models with Foundation Models

Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their performance. The traditional solution is fine-tuning, but this undermines the key advantage of pretrained models, which is their ability to be used out-of-the-box. We propose ROBOSHOT, a…

Oct 20, 2023

D. Adila, et al.

Learn more about Zero-Shot Robustification of Zero-Shot Models with Foundation Models
Skill-It! A Data-Driven Skills Framework for Understanding and Training Language Models
The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good downstream model performance across tasks. We develop a new framework based on a simple hypothesis: just as humans acquire interdependent skills in a deliberate order, language models also follow a natural order when learning a set of skills from their training data. If such an order exists, it can be utilized for improved understanding of LMs and for data-efficient training. Using this intuition, our framework formalizes the notion of...
Research Paper
Skill-It! A Data-Driven Skills Framework for Understanding and Training Language Models

The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good downstream model performance across tasks. We develop a new framework based on a simple hypothesis: just as humans acquire interdependent skills in a deliberate order, language models also follow…

Oct 20, 2023

MF Chen, et al.

Learn more about Skill-It! A Data-Driven Skills Framework for Understanding and Training Language Models
1 24 25 26 65
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.