Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho · 2024 · Stanford RegLab and the Institute for Human-Centered AI; arXiv 2405.20362
Legal AI is an unusually evidence-poor field. Almost every accuracy claim a firm encounters originates with the company making the sale, tested on questions that company chose, using a definition of error that company wrote. This study is the exception, and that alone would justify its place here. Researchers at Stanford built their own question set, ran it against the leading commercial legal research products, and reported what came back — finding hallucination rates between roughly 17% and 33% depending on the product. On the better-performing system that is more than one query in six returning misleading or false information; on the weaker one, close to a third of responses.
But the number is not the important part, because numbers move between releases and vendors have since published better ones. The important part is structural. Both major vendors had marketed retrieval-augmented architectures on the promise that grounding the model in a real corpus prevents hallucination, and the study examined that claim directly and found it overstated. Retrieval does substantially reduce citations to authority that does not exist — that much is real and it matters. What it does not touch is the harder failure: a real case retrieved and then characterised wrongly, a holding overstated, a synthesis no retrieved document supports. Those errors arrive attached to citations that check out, which makes them considerably harder for a reviewer to catch than a fabricated case would be. A firm that hears 'grounded' and concludes 'verification is now optional' has drawn precisely the wrong inference from a genuine improvement.
Vendors disputed aspects of the methodology, particularly the composition of the question set and what was scored as a hallucination, and have since published figures showing substantial improvement — citation hallucination in low single digits has been reported for at least one product, against materially higher figures the year before. Both positions can hold: a study measures a system at a moment, and these systems improve. What survives the dispute entirely is the architectural finding, because no party claims the rate is zero, and it is only zero that would let a firm stop checking.