What does retrieval-augmented generation actually do to an AI legal research tool?
Retrieval-augmented generation puts a search step in front of the language model, so the model writes its answer from documents the system just pulled rather than from what it absorbed in training. The paper by Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho diagrams the pipeline in two stages: retrieval, where the query is embedded and a retrieval system returns candidate documents, and generation, where those documents are handed to the model to produce the response.
The reason this matters to a lawyer is narrow and real. A model working from memory can produce a case name, a reporter volume and a year that have never existed together, because it is predicting plausible text. A model working from retrieved documents is far less likely to do that, because the case in front of it is an actual case.
What the architecture does not do is guarantee that the retrieved case is the right case, that the passage relied on means what the answer says, or that the tool will decline when the database holds no answer. The authors note that each subsidiary step in the pipeline can fail on its own. The Institute’s Grounding & retrieval page walks through those failure points.
What did the Stanford study of AI legal research tools actually measure, and what did it find?
It measured how often three commercial legal research products produced false or misleading answers, and found rates between 17% and 33%. The study is Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, by Varun Magesh, Faiz Surani, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho of Stanford University with Matthew Dahl of Yale University. It was posted to arXiv on 30 May 2024, updated on 3 June 2024 to add Westlaw AI-Assisted Research, and published in the Journal of Empirical Legal Studies, volume 22, at pages 216 to 242, in 2025.
The method was a manually constructed, preregistered set of 202 legal queries run against Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI, with GPT-4 as a comparison. Queries fell into four families: general research questions, jurisdiction-specific or time-specific questions, false-premise questions, and factual recall questions. Each response was coded for correctness — correct, incorrect, or a refusal — and correct responses were then coded for groundedness.
The reported results differ sharply between products. The authors wrote that Lexis+ AI answered 65% of queries accurately, that Westlaw AI-Assisted Research was accurate 42% of the time while hallucinating nearly twice as often as the other legal tools tested, and that Ask Practical Law AI gave incomplete answers — refusals or ungrounded responses — on more than 60% of queries. All three hallucinated less than GPT-4. None of them hallucinated at zero, which is the only number that would let a firm skip verification. The Accuracy evidence page collects what each of these figures does and does not cover.
What is the difference between a hallucinated citation and a real case cited for a proposition it does not support?
A hallucinated citation points to something that does not exist; a misgrounded citation points to something real that does not support the sentence attached to it. The Magesh and Surani paper draws the line formally. A response is grounded when its key factual propositions make valid references to relevant legal documents, ungrounded when key propositions are not cited at all, and misgrounded when key propositions “are cited but misinterpret the source or reference an inapplicable source.”
That third category is the one retrieval does not solve, and it is the one a citation check will not catch. Running the cite through a validation service confirms the case exists and is good law. It says nothing about whether the case holds what the paragraph claims. The verification step that catches a misgrounded answer is reading the cited authority, which is the step a firm is most tempted to drop once the tool starts sounding fluent.
The authors also found that the tools were vulnerable to false-premise questions, where the query itself assumes something untrue. Matthew Dahl, Varun Magesh, Mirac Suzgun and Daniel E. Ho had documented the same pattern in general-purpose models in Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis, volume 16, issue 1 (2024), where models tended to accept a false premise rather than correct it. A tool that answers the question you asked, rather than telling you the question is wrong, is a specific hazard for a lawyer who is already partway to a theory.
How did the vendors respond to the findings?
Thomson Reuters disputed the accuracy figure publicly, and the exchange is worth reading because of what each side was measuring. On 10 June 2024, Mike Dahn, head of Westlaw Product Management at Thomson Reuters, wrote that the company was “very supportive of efforts to test and benchmark solutions like this” but was “quite surprised when we saw the claims of significant issues with hallucinations with AI-Assisted Research,” and said that internal testing showed “an accuracy rate of approximately 90% based on how our customers use it.”
Dahn described that internal testing as hundreds of real-world legal research questions, with two lawyers grading each result and a third, more senior lawyer resolving disagreements, and he suggested the researchers found higher inaccuracy because the study “included question types we very rarely or never see in AI-Assisted Research.” He also wrote that the product uses large language models and “can occasionally produce inaccuracies,” and described it as an accelerant rather than a replacement for thorough research. Every number in that paragraph is vendor-reported and drawn from the company’s own testing; no independent replication of it has been published.
The sequence around access is also part of the record. Bob Ambrogi reported at LawSites on 4 June 2024 that the original version of the study had omitted Westlaw AI-Assisted Research because Thomson Reuters had denied the researchers access, that the company subsequently provided access, and that the updated study then put Westlaw at 42% accuracy and a 33% hallucination rate against Lexis+ AI at 65% and 17%. That is the shape of most disputes in this market: the vendor and the researcher are not testing the same questions, and only one of them publishes the question set.
Has anyone measured legal research tools since then?
Yes, though the later public benchmark has a participation problem that a reader has to weigh. Vals AI published a legal-research-specific edition of the Vals Legal AI Report, dated 14 October 2025, following its first report of 27 February 2025. The legal research study put 200 questions across ten question types to each participant, with one question disregarded for a formulation error, and had responses graded blind by a group of lawyers and law librarians, each response seen by at least two assessors and zero scores escalated to a third.
The participants were Alexi, Counsel Stack and Midpage on the legal side and OpenAI’s ChatGPT as a generalist comparison, against a baseline of lawyers supplied by a United States law firm that partnered with Vals AI, who answered the same prompts in writing within a two-week period and were allowed every research resource at their disposal except generative AI tools. The report scored on accuracy weighted at 50%, authoritativeness of sources at 40% and appropriateness at 10%, and reported the participating legal products between 78% and 81% on accuracy, ChatGPT at 80%, and the lawyer baseline at 69%.
Two caveats travel with those numbers. Participation was voluntary, and the participant list contains no product from LexisNexis or Thomson Reuters, so this is not a measurement of the products most firms actually license. And the report itself notes that posing a research question as a single zero-shot prompt “is not necessarily true-to-life,” with no follow-up prompting and no workflow features in the test. A benchmark that a vendor chose to enter is evidence about that vendor on that day, not a ranking of the market. The Institute does not rank products; Evaluating tools covers testing a tool against your own matters instead.
What verification does the lawyer still owe regardless of which tool produced the answer?
Reading the cited authority, because that is the only step that catches the error retrieval leaves behind. ABA Formal Opinion 512, issued by the ABA Standing Committee on Ethics and Professional Responsibility on 29 July 2024 and titled Generative Artificial Intelligence Tools, states that generative tools “lack the ability to understand the meaning of the text they generate or evaluate its context” and that “lawyers’ uncritical reliance on content created by a GAI tool can result in inaccurate legal advice to clients or misleading representations to courts and third parties.”
The opinion does not set a fixed amount of checking. It says the appropriate amount of independent verification or review “will necessarily depend on the GAI tool and the specific task that it performs,” and gives the example of a lawyer who has already tested a summarization tool against a manually reviewed subset and found it accurate. Formal Opinion 512 is the ABA’s guidance, not a rule in any jurisdiction; rules of professional conduct are adopted state by state and diverge, and what your jurisdiction requires is a question for counsel there.
What follows practically is a workflow question rather than an ethics question. A verification step that only confirms cases exist is calibrated to the 2023 failure mode. The 2024 through 2026 evidence says the surviving failure is a real case pressed into service for a proposition it does not carry, and the only thing that finds it is a person reading the case. The Institute’s Citation verification and Research workflow pages cover how firms build that step so it survives a deadline, and how often legal AI still gets it wrong takes up the error-rate question across tasks beyond research.