Resource library


Enterprises must evaluate LLM performance for production deployment. Custom, automated eval + data slices present the best path to production.
In this webinar, presented by Snorkel and Numbers Station, we’ll explain how to elicit data-driven insights backed by domain-knowledge using a multi-agent architecture and a fine-tuned LLM and/or RAG pipeline.


Meta’s Llama 3.1 405B, rivals GPT-4o in benchmarks, offering powerful AI capabilities. Despite high costs, it can enhance LLM adoption through fine-tuning, distillation, and as an AI judge.


Meta released Llama 3 405B today, signaling a new era of open source AI. The model is ready to use on Snorkel Flow.


High-performing AI systems require more than a well-designed model. They also require properly constructed training and testing data.


We need more labeled data than ever, so we have explored weak supervision for non-categorical applications—with notable results.


In this webinar, Vincent Chen, Product Director at Snorkel AI, will discuss the importance of LLM evaluation, highlight common challenges and approaches, explain core concepts such as slices and quality models, and demonstrate Snorkel AI’s approach to LLM evaluation.


To tackle generative AI use cases, Snorkel AI + AWS launched an accelerator program to address the biggest blocker: unstructured data.


AI alignment ensures that AI systems align with human values, ethics, and policies. Here’s a primer on how developers can build safer AI.












