home  /  library  /  hallucination-free-legal-research-tools
Independent studies

Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools

Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho · 2024 · Stanford RegLab and the Institute for Human-Centered AI; arXiv 2405.20362

Legal researchVerificationChoosing toolsAnyone deciding whether to rely on an AI legal research tool
Why it is on the shelf. The single most important document on this shelf, and the only place a firm can find independently measured error rates for the commercial legal research tools it is being sold. Its finding on retrieval architecture is the one that should change how a firm designs its verification step.

The Institute's reading

Legal AI is an unusually evidence-poor field. Almost every accuracy claim a firm encounters originates with the company making the sale, tested on questions that company chose, using a definition of error that company wrote. This study is the exception, and that alone would justify its place here. Researchers at Stanford built their own question set, ran it against the leading commercial legal research products, and reported what came back — finding hallucination rates between roughly 17% and 33% depending on the product. On the better-performing system that is more than one query in six returning misleading or false information; on the weaker one, close to a third of responses.

But the number is not the important part, because numbers move between releases and vendors have since published better ones. The important part is structural. Both major vendors had marketed retrieval-augmented architectures on the promise that grounding the model in a real corpus prevents hallucination, and the study examined that claim directly and found it overstated. Retrieval does substantially reduce citations to authority that does not exist — that much is real and it matters. What it does not touch is the harder failure: a real case retrieved and then characterised wrongly, a holding overstated, a synthesis no retrieved document supports. Those errors arrive attached to citations that check out, which makes them considerably harder for a reviewer to catch than a fabricated case would be. A firm that hears 'grounded' and concludes 'verification is now optional' has drawn precisely the wrong inference from a genuine improvement.

Key propositions

  • Leading commercial legal research tools hallucinated in roughly 17% to 33% of responses at the time of testing.
  • Retrieval-augmented generation reduces fabricated citations substantially but does not eliminate hallucination.
  • Vendor claims that these architectures prevent hallucination were found to be overstated.
  • The surviving errors — mischaracterised holdings, unsupported synthesis — are harder to detect than fabricated ones.
  • Purpose-built legal tools nonetheless performed materially better than general-purpose models on legal queries.

In practice

  • Verification to primary sources cannot be dropped on the strength of a grounding claim, whatever the product.
  • Review effort should be aimed at whether authority says what the answer claims — not only at whether it exists.
  • Any accuracy figure a firm relies on should carry the date it was measured, because these systems change between releases.

Where authorities disagree

Vendors disputed aspects of the methodology, particularly the composition of the question set and what was scored as a hallucination, and have since published figures showing substantial improvement — citation hallucination in low single digits has been reported for at least one product, against materially higher figures the year before. Both positions can hold: a study measures a system at a moment, and these systems improve. What survives the dispute entirely is the architectural finding, because no party claims the rate is zero, and it is only zero that would let a firm stop checking.

AI adoption conciergeorientation · not legal or ethics advice
Happy to dig into it. What would you like to pressure-test from Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools: one of its propositions, how it applies to your situation, or where it disagrees with the rest of the shelf?