resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers and more.
Image for Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark
Blog

Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark

Announcing a $3M commitment to launch Open Benchmarks Grants
September 30, 2025
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory
Blog

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

Announcing a $3M commitment to launch Open Benchmarks Grants
March 31, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
of
Type: All Types
Sort: Newest
BERT models: Google’s NLP for the enterprise
Blog
BERT models: Google’s NLP for the enterprise

LLMs have claimed the spotlight since the debut of ChatGPT, but BERT models quietly handle most enterprise production NLP tasks.

Dec 27, 2023
Learn more about BERT models: Google’s NLP for the enterprise
First cohort of Snorkel GenAI customers sees gains up to 54 points
Blog
First cohort of Snorkel GenAI customers sees gains up to 54 points

In its first six months, Snorkel Foundry collaborated on high-value projects with notable companies and produced impressive results.

Dec 20, 2023
Learn more about First cohort of Snorkel GenAI customers sees gains up to 54 points
How to tackle advanced classification challenges using Snorkel Flow
Blog
How to tackle advanced classification challenges using Snorkel Flow

When done right, advanced classification applications cultivate business value and automation, unlock new business lines, and reduce costs.

Dec 14, 2023
Learn more about How to tackle advanced classification challenges using Snorkel Flow
How to scale chatbot development with Google Dialogflow and Snorkel Flow
Blog
How to scale chatbot development with Google Dialogflow and Snorkel Flow

A brief guide on how financial institutions could use Google Dialogflow with Snorkel Flow to build better chatbots for retail banking

Dec 12, 2023
Learn more about How to scale chatbot development with Google Dialogflow and Snorkel Flow
Foundation Models Can Robustify Themselves, For Free
Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their performance. The traditional solution is fine-tuning, but this undermines the key advantage of pretrained models, which is their ability to be used out-of-the-box. We propose ROBOSHOT, a method that improves the robustness of pretrained model embeddings in a fully zero-shot fashion. First, we use language models (LMs) to obtain useful insights from task descriptions. These insights are embedded and used to remove harmful and boost useful components...
Research Paper
Foundation Models Can Robustify Themselves, For Free

Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their performance. The traditional solution is fine-tuning, but this undermines the key advantage of pretrained models, which is their ability to be used out-of-the-box. We propose ROBOSHOT, a…

Dec 12, 2023

D. Adila, et al.

Learn more about Foundation Models Can Robustify Themselves, For Free
Train ‘n Trade: Foundations of Parameter Markets
Organizations typically train large models individually. This is costly and time-consuming, particularly for large-scale foundation models. Such vertical production is known to be suboptimal. Inspired by this economic insight, we ask whether it is possible to leverage others’ expertise by trading the constituent parts in models, i.e., sets of weights, as if they were market commodities. While recent advances in aligning and interpolating models suggest that doing so may be possible, a number of fundamental questions must be answered to create viable parameter markets. In this work, we address these basic questions, propose a framework containing the infrastructure necessary for...
Research Paper
Train ‘n Trade: Foundations of Parameter Markets

Organizations typically train large models individually. This is costly and time-consuming, particularly for large-scale foundation models. Such vertical production is known to be suboptimal. Inspired by this economic insight, we ask whether it is possible to leverage others’ expertise by trading the constituent parts in models, i.e., sets of weights, as if they were market commodities. While recent advances in…

Dec 07, 2023

TH. Huang, et al.

Learn more about Train ‘n Trade: Foundations of Parameter Markets
How predictive AI + generative AI build amazing document understanding
Blog
How predictive AI + generative AI build amazing document understanding

A proof-of-concept project that combines predictive AI + generative AI to minimize LLM’s risks while keeping their advantages.

Dec 05, 2023
Learn more about How predictive AI + generative AI build amazing document understanding
The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
Compressing large language models (LLMs), often consisting of billions of parameters, provides faster inference, smaller memory footprints, and enables local deployment. Two standard compression techniques are pruning and quantization, with the former eliminating redundant connections in model layers and the latter representing model parameters with fewer bits. The key tradeoff is between the degree of compression and the impact on the quality of the compressed model. Existing research on LLM compression primarily focuses on performance in terms of general metrics like perplexity or downstream task accuracy. More fine-grained metrics, such as those measuring parametric knowledge, remain significantly underexplored. To help...
Research Paper
The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models

Compressing large language models (LLMs), often consisting of billions of parameters, provides faster inference, smaller memory footprints, and enables local deployment. Two standard compression techniques are pruning and quantization, with the former eliminating redundant connections in model layers and the latter representing model parameters with fewer bits. The key tradeoff is between the degree of compression and the impact on…

Dec 02, 2023

SSS. Namburi, et al.

Learn more about The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
How to fine-tune Llama 2 in Snorkel Flow
Blog
How to fine-tune Llama 2 in Snorkel Flow

Data scientists can fine-tune Llama 2 to adapt it to specific tasks. The Snorkel Flow data development platform makes it easy to do so.

Nov 28, 2023
Learn more about How to fine-tune Llama 2 in Snorkel Flow
1 22 23 24 65
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.