resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers and more.
Image for Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark
Blog

Evaluating Coding Agent Capabilities with Terminal-Bench: Snorkel’s Role in Building the Next Generation Benchmark

Announcing a $3M commitment to launch Open Benchmarks Grants
September 30, 2025
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory
Blog

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory

Announcing a $3M commitment to launch Open Benchmarks Grants
March 31, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
of
Type: All Types
Sort: Newest
AI data development: a guide for data science projects
Blog
AI data development: a guide for data science projects

What is AI data development? AI data development includes any action taken to convert raw information into a format useful to AI.

Nov 13, 2024
Learn more about AI data development: a guide for data science projects
Webinar
GenAI evaluations that identify weaknesses and actionable insights

In this webinar, we’ll explain how enterprise AI/ML teams use Snorkel Flow’s evaluation framework to measure GenAI performance based on SME guidelines and feedback, identify areas of improvement and take corrective action to reach production accuracy requirements.

Nov 13, 2024
Snorkel Team
Learn more about GenAI evaluations that identify weaknesses and actionable insights
Webinar
Building specialized LLMs with Bedrock + SageMaker

In this demo, learn how Snorkel AI powers efficient data labeling and management for LLM fine-tuning, natively integrating with Amazon Bedrock and SageMaker for seamless, end-to-end development of specialized models.

Oct 29, 2024
Snorkel Team
Learn more about Building specialized LLMs with Bedrock + SageMaker
SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies
Blog
SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies

Discover highlights of Snorkel AI’s first annual SnorkelCon user conference. Explore Snorkel’s programmatic AI data development achievements.

Oct 22, 2024
Learn more about SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies
Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows
Blog
Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows

Snorkel AI has made building production-ready, high-value enterprise AI applications faster and easier than ever. The 2024.R3 update to our Snorkel Flow AI data development platform streamlines data-centric workflows, from easier-than-ever generative AI evaluation to multi-schema annotation.

Oct 09, 2024
Learn more about Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows
Explore the new GenAI Evaluation Suite: Snorkel 2024.R3
Blog
Explore the new GenAI Evaluation Suite: Snorkel 2024.R3

We aim to help our customers get GenAI into production. In our 2024.R3 release, we’ve delivered some exciting GenAI evaluation results.

Oct 09, 2024
Learn more about Explore the new GenAI Evaluation Suite: Snorkel 2024.R3
New NLP features in Snorkel Flow 2024.R3
Blog
New NLP features in Snorkel Flow 2024.R3

Discover new NLP features in Snorkel Flow\’s 2024.R3 release, including named entity recognition for PDFs + advanced sequence tagging tools.

Oct 09, 2024
Learn more about New NLP features in Snorkel Flow 2024.R3
Enterprise data compliance and security review: Snorkel Flow 2024.R3
Blog
Enterprise data compliance and security review: Snorkel Flow 2024.R3

Discover the latest enterprise readiness features for Snorkel Flow. Configure safeguards for data compliance and security.

Oct 09, 2024
Learn more about Enterprise data compliance and security review: Snorkel Flow 2024.R3
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
Existing ML benchmarks lack the depth and diversity of annotations needed for evaluating models on business process management (BPM) tasks. BPM is the practice of documenting, measuring, improving, and automating enterprise workflows. However, research has focused almost exclusively on one task– full end-to-end automation using agents based on multimodal foundation models (FMs) like GPT-4. This focus on automation ignores the reality of how most BPM tools are applied today– simply documenting the relevant workflow takes 60% of the time of the typical process optimization project. To address this gap we present WONDERBREAD, the first benchmark for evaluating multimodal FMs on...
Research Paper
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks

Existing ML benchmarks lack the depth and diversity of annotations needed for evaluating models on business process management (BPM) tasks. BPM is the practice of documenting, measuring, improving, and automating enterprise workflows. However, research has focused almost exclusively on one task– full end-to-end automation using agents based on multimodal foundation models (FMs) like GPT-4. This focus on automation ignores the…

Oct 01, 2024

Michael Wornow Avanika Narayan Ben Viggiano Ishan S. Khare Tathagat Verma Tibor Thompson Miguel Angel Fuentes Hernandez Sudharsan Sundar Chloe Trujillo Krrish Chawla Rongfei Lu Justin Shen Divya Nagaraj Joshua Martinez Vardhan Agrawal Althea Hudson Nigam H. Shah Christopher Ré Stanford University

Learn more about WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
1 9 10 11 65
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.