Resource library
Introduction: With the advent of consumer-facing devices that can render atrial fibrillation (AF) pre-diagnosis, medical wearables now have the potential to affect diagnosis rates and medical care. Post-market surveillance is necessary to understand the impact of wearables on patient outcomes and health care utilization, but is hindered by the lack of codified terms in EHR that capture wearable use. Research…


LLMs have a broad but shallow knowledge, but fall short on specialized tasks. For best performance, enterprises must fine tune their LLMs.


The 2023.R3 Snorkel Flow release is packed with improvements that amplify user experience, streamline workflows, and enhance performance, ensuring our users derive unparalleled value from our platform.


The Biden administration issued an executive order that creates new AI standards and challenges. AI data development can help.


The Snorkel AI team will present 18 research papers and talks at the 2023 Neural Information Processing Systems (NeurIPS) conference from December 10-16. The Snorkel papers cover a broad range of topics including fairness, semi-supervised learning, large language models (LLMs), and domain-specific models. Snorkel AI is proud of its roots in the research community and endeavors to remain at the forefront…


Distillation techniques allow enterprises to access the full predictive power of large language models at a tiny fraction of their cost.


Snorkel AI’s Enterprise LLM Virtual Summit drew 1,000 attendees with speakers from Contextual AI, Google, Meta, Stanford, and Together AI.


Zero-shot inference is a powerful paradigm that enables the use of large pretrained models for downstream classification tasks without further training. However, these models are vulnerable to inherited biases that can impact their performance. The traditional solution is fine-tuning, but this undermines the key advantage of pretrained models, which is their ability to be used out-of-the-box. We propose ROBOSHOT, a…


The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good downstream model performance across tasks. We develop a new framework based on a simple hypothesis: just as humans acquire interdependent skills in a deliberate order, language models also follow…












