A Harvey benchmark, built with data partner Snorkel AI, for hard agentic legal research problems, testing whether AI can search, reason, and cite case law the way a practicing lawyer would, end to end.
BigLaw Bench: Research (BLB: Research) is a benchmark focused on hard agentic legal research problems. Working with its data partner Snorkel AI, a leader in creating complex expert data for frontier AI, Harvey identified a series of US case law research problems that leading models currently cannot solve, even when given search tools like web search. It is the second of Harvey’s major BigLaw Bench expansions this quarter.
15
legal practice
areas
covered
<60%
task completion below
which answers become
unhelpful
END-TO-END
search + grounded,
cited response
in one task
BLB: Research serves two purposes: to identify how foundation models are improving on research tasks and the failure modes that still cause unsatisfactory answers, and to investigate ways to improve AI-based legal research through agentic systems and access to non-public data, not just raw model capability. By revealing both the best model and where better infrastructure is needed, it points the way to deeper, more accurate legal research.
Building a holistic search benchmark allows us to measure performance in the way that tracks how research capabilities are being developed and how our customers expect them to be delivered.
The reasoning: search is now the primary means of grounding model responses across the industry. As models are increasingly trained to search rather than rely on training data for up-to-date knowledge, benchmarking them without search undersells what they can do — and pin-cited sources have become table stakes for research-driven legal work.
Models are already very good at search — they can write boolean and web queries and refine their research. So the goal of BLB: Research was to validate those baseline capabilities and then find where models still cannot turn search into meaningful research outcomes for lawyers. The challenge was building a benchmark that is hard and realistic: esoteric tasks are easy to fail but useless if success wouldn’t actually help a legal professional.
Rather than predict what should be hard, Harvey found the frontier of model capability empirically. First it defined failure on a structured rubric — how poorly a model must perform before its answer is unhelpful. Many partly correct answers are still useful (slightly-off analysis with good citations still gives a research head start). In general, answers become unhelpful once a model completes less than 60% of the required task criteria — typically by missing critical reasoning junctures, taking wrong turns, or staying too surface-level.
Representative tasks Harvey uses to benchmark model research capabilities.
| Practice area | Objective | Description |
|---|---|---|
| Corporate | Assess earn-out manipulation claims after an asset sale | A strategic buyer acquires a Delaware company’s assets with a multi-year earn-out, then consolidates the business, reallocates overhead, shifts pricing, and reassigns salespeople. Assess whether the integration decisions breach the Asset Purchase Agreement. |
| Securities Litigation | Evaluate securities fraud claims for prototype misrepresentations | A hydrogen fuel-cell truck company’s CEO made statements about functional prototypes and partnerships that internal documents contradict. Evaluate material misrepresentation, scienter pleading, and PSLRA safe harbor after a 45% stock drop. |
| Privacy & Cybersecurity | Analyze defenses to a nationwide class action from a data breach | A fintech with 650,000 affected users faces a putative nationwide class action after a breach from a compromised contractor API credential. Evaluate Article III standing, choice-of-law barriers to certification, and Rule 12(b)(6) strategies. |
Other practice areas in the benchmark include Intellectual Property, Commercial Litigation, Constitutional Law, Regulatory, Employment & Labor, Health Law & Life Sciences, Tax, Tort, Real Property, Media and Technology, Immigration, and Family Law.
Sifting through documents to find answers, narratives, and trends is one of the most common tasks in legal practice. The same search capabilities that make models good at case law research also underpin searching EDGAR, an investigation corpus, a deal room, or a firm’s own internal knowledge. BLB: Research will build toward measuring all of these, tracking innovations in model capability, unique data sources, and the tools that connect them, and converting them into better research results for legal teams.
Snorkel AI served as Harvey’s data partner for BigLaw Bench: Research, bringing its expertise in creating complex expert data for frontier AI to identify the US case law research problems that today’s leading models cannot yet solve. The collaboration reflects Snorkel’s broader work in data-centric AI research and expert-grounded evaluation for high-stakes domains.
BigLaw Bench: Research is part of Snorkel’s broader work on AI agent evaluation. Explore related benchmarks: JudgmentBench (another Harvey + Snorkel legal AI collaboration), Agents’ Last Exam, Terminal-Bench Science, and Terminal-Bench 2.0.