Resource library


Snorkel and affiliated academic labs have been hard at work reducing how computationally expensive large language models are.


Large language models (LLMs) have been recently leveraged as training data generators for various natural language processing (NLP) tasks. While previous research has explored different approaches to training models using generated data, they generally rely on simple class-conditional prompts, which may limit the diversity of the generated data and inherit systematic biases of LLM. Thus, we investigate training data generation…


Claypot AI CEO Chip Huyen presented “Platform for Real-Time Machine Learning” at Snorkel AI’s Future of Data-Centric AI 2022.


Enterprises—especially the world’s largest—are excited to use large language models, but they want to fine-tune them on proprietary data.


Jacomo Corbo and Bryan Richardson with QuantumBlack present “Automating Data Quality Remediation With AI” at The Future of Data-Centric AI.


Stefano Lindt presents “Leveraging NLP to Extract Value From Business Data” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022.


Peter Davio, CTO at Black Swan Data, presented “Petabyte-Level Learning” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022.


Grammarly’s Timo Mertens presents “Toward Superhuman Communication Assistance” at Snorkel AI’s The Future of Data-Centric AI Summit in 2022.


The Future of Data-Centric AI showcased customer to successes, took a deep look at Snorkel Flow, and announced two new solutions.












