home  /  insights  /  what-rag-changes-about-legal-research
Legal Research with AI

Does retrieval-augmented generation fix AI legal research?

Retrieval cuts invented authority. It does not cut the other failure — a real case, correctly cited, standing for something it never held. Six researchers at Stanford and Yale measured both, and Thomson Reuters disputed the result.

September 15, 2026 · 10 min read

The short answer

It fixes part of the problem and leaves the harder part untouched. Retrieval-augmented generation makes a model answer from documents pulled out of a database rather than from memory alone, which substantially reduces wholly invented citations. In the first preregistered independent test of the commercial legal research tools — Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, posted 30 May 2024 and published in the Journal of Empirical Legal Studies, volume 22, pages 216 to 242, in 2025 — the tools built on that architecture still produced hallucinated answers between 17% and 33% of the time. The error that survives retrieval is not a fake case. It is a real case, correctly cited, that does not say what the answer says it says.

What this article establishes

  • The Stanford and Yale team of Magesh, Surani, Dahl, Suzgun, Manning and Ho ran 202 preregistered queries against Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI and GPT-4, and reported hallucination rates of 17% to 33% for the three legal tools.
  • The same paper quotes the marketing it was testing: LexisNexis on “100% hallucination-free linked legal citations,” Thomson Reuters on “checks and balances that ensure our answers are grounded in good law,” and Casetext that CoCounsel “does not make up facts.” The authors concluded those claims were overstated.
  • The paper’s own typology separates “incorrect” from “misgrounded” — a misgrounded response states the law accurately but cites a source that misinterprets it or does not apply. A citation checker catches neither, because the case is real and the reporter cite is right.
  • Thomson Reuters disputed the finding. On 10 June 2024 Mike Dahn, head of Westlaw Product Management, wrote that internal testing showed “an accuracy rate of approximately 90% based on how our customers use it,” graded by two lawyers per result with a third resolving disagreements. That figure is vendor-reported and was not independently replicated.
  • ABA Formal Opinion 512, issued 29 July 2024 by the Standing Committee on Ethics and Professional Responsibility, states that “lawyers’ uncritical reliance on content created by a GAI tool can result in inaccurate legal advice to clients or misleading representations to courts,” and ties the amount of verification required to the tool and the task.

What does retrieval-augmented generation actually do to an AI legal research tool?

Retrieval-augmented generation puts a search step in front of the language model, so the model writes its answer from documents the system just pulled rather than from what it absorbed in training. The paper by Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho diagrams the pipeline in two stages: retrieval, where the query is embedded and a retrieval system returns candidate documents, and generation, where those documents are handed to the model to produce the response.

The reason this matters to a lawyer is narrow and real. A model working from memory can produce a case name, a reporter volume and a year that have never existed together, because it is predicting plausible text. A model working from retrieved documents is far less likely to do that, because the case in front of it is an actual case.

What the architecture does not do is guarantee that the retrieved case is the right case, that the passage relied on means what the answer says, or that the tool will decline when the database holds no answer. The authors note that each subsidiary step in the pipeline can fail on its own. The Institute’s Grounding & retrieval page walks through those failure points.

It measured how often three commercial legal research products produced false or misleading answers, and found rates between 17% and 33%. The study is Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, by Varun Magesh, Faiz Surani, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho of Stanford University with Matthew Dahl of Yale University. It was posted to arXiv on 30 May 2024, updated on 3 June 2024 to add Westlaw AI-Assisted Research, and published in the Journal of Empirical Legal Studies, volume 22, at pages 216 to 242, in 2025.

The method was a manually constructed, preregistered set of 202 legal queries run against Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI, with GPT-4 as a comparison. Queries fell into four families: general research questions, jurisdiction-specific or time-specific questions, false-premise questions, and factual recall questions. Each response was coded for correctness — correct, incorrect, or a refusal — and correct responses were then coded for groundedness.

The reported results differ sharply between products. The authors wrote that Lexis+ AI answered 65% of queries accurately, that Westlaw AI-Assisted Research was accurate 42% of the time while hallucinating nearly twice as often as the other legal tools tested, and that Ask Practical Law AI gave incomplete answers — refusals or ungrounded responses — on more than 60% of queries. All three hallucinated less than GPT-4. None of them hallucinated at zero, which is the only number that would let a firm skip verification. The Accuracy evidence page collects what each of these figures does and does not cover.

What is the difference between a hallucinated citation and a real case cited for a proposition it does not support?

A hallucinated citation points to something that does not exist; a misgrounded citation points to something real that does not support the sentence attached to it. The Magesh and Surani paper draws the line formally. A response is grounded when its key factual propositions make valid references to relevant legal documents, ungrounded when key propositions are not cited at all, and misgrounded when key propositions “are cited but misinterpret the source or reference an inapplicable source.”

That third category is the one retrieval does not solve, and it is the one a citation check will not catch. Running the cite through a validation service confirms the case exists and is good law. It says nothing about whether the case holds what the paragraph claims. The verification step that catches a misgrounded answer is reading the cited authority, which is the step a firm is most tempted to drop once the tool starts sounding fluent.

The authors also found that the tools were vulnerable to false-premise questions, where the query itself assumes something untrue. Matthew Dahl, Varun Magesh, Mirac Suzgun and Daniel E. Ho had documented the same pattern in general-purpose models in Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis, volume 16, issue 1 (2024), where models tended to accept a false premise rather than correct it. A tool that answers the question you asked, rather than telling you the question is wrong, is a specific hazard for a lawyer who is already partway to a theory.

How did the vendors respond to the findings?

Thomson Reuters disputed the accuracy figure publicly, and the exchange is worth reading because of what each side was measuring. On 10 June 2024, Mike Dahn, head of Westlaw Product Management at Thomson Reuters, wrote that the company was “very supportive of efforts to test and benchmark solutions like this” but was “quite surprised when we saw the claims of significant issues with hallucinations with AI-Assisted Research,” and said that internal testing showed “an accuracy rate of approximately 90% based on how our customers use it.”

Dahn described that internal testing as hundreds of real-world legal research questions, with two lawyers grading each result and a third, more senior lawyer resolving disagreements, and he suggested the researchers found higher inaccuracy because the study “included question types we very rarely or never see in AI-Assisted Research.” He also wrote that the product uses large language models and “can occasionally produce inaccuracies,” and described it as an accelerant rather than a replacement for thorough research. Every number in that paragraph is vendor-reported and drawn from the company’s own testing; no independent replication of it has been published.

The sequence around access is also part of the record. Bob Ambrogi reported at LawSites on 4 June 2024 that the original version of the study had omitted Westlaw AI-Assisted Research because Thomson Reuters had denied the researchers access, that the company subsequently provided access, and that the updated study then put Westlaw at 42% accuracy and a 33% hallucination rate against Lexis+ AI at 65% and 17%. That is the shape of most disputes in this market: the vendor and the researcher are not testing the same questions, and only one of them publishes the question set.

Yes, though the later public benchmark has a participation problem that a reader has to weigh. Vals AI published a legal-research-specific edition of the Vals Legal AI Report, dated 14 October 2025, following its first report of 27 February 2025. The legal research study put 200 questions across ten question types to each participant, with one question disregarded for a formulation error, and had responses graded blind by a group of lawyers and law librarians, each response seen by at least two assessors and zero scores escalated to a third.

The participants were Alexi, Counsel Stack and Midpage on the legal side and OpenAI’s ChatGPT as a generalist comparison, against a baseline of lawyers supplied by a United States law firm that partnered with Vals AI, who answered the same prompts in writing within a two-week period and were allowed every research resource at their disposal except generative AI tools. The report scored on accuracy weighted at 50%, authoritativeness of sources at 40% and appropriateness at 10%, and reported the participating legal products between 78% and 81% on accuracy, ChatGPT at 80%, and the lawyer baseline at 69%.

Two caveats travel with those numbers. Participation was voluntary, and the participant list contains no product from LexisNexis or Thomson Reuters, so this is not a measurement of the products most firms actually license. And the report itself notes that posing a research question as a single zero-shot prompt “is not necessarily true-to-life,” with no follow-up prompting and no workflow features in the test. A benchmark that a vendor chose to enter is evidence about that vendor on that day, not a ranking of the market. The Institute does not rank products; Evaluating tools covers testing a tool against your own matters instead.

What verification does the lawyer still owe regardless of which tool produced the answer?

Reading the cited authority, because that is the only step that catches the error retrieval leaves behind. ABA Formal Opinion 512, issued by the ABA Standing Committee on Ethics and Professional Responsibility on 29 July 2024 and titled Generative Artificial Intelligence Tools, states that generative tools “lack the ability to understand the meaning of the text they generate or evaluate its context” and that “lawyers’ uncritical reliance on content created by a GAI tool can result in inaccurate legal advice to clients or misleading representations to courts and third parties.”

The opinion does not set a fixed amount of checking. It says the appropriate amount of independent verification or review “will necessarily depend on the GAI tool and the specific task that it performs,” and gives the example of a lawyer who has already tested a summarization tool against a manually reviewed subset and found it accurate. Formal Opinion 512 is the ABA’s guidance, not a rule in any jurisdiction; rules of professional conduct are adopted state by state and diverge, and what your jurisdiction requires is a question for counsel there.

What follows practically is a workflow question rather than an ethics question. A verification step that only confirms cases exist is calibrated to the 2023 failure mode. The 2024 through 2026 evidence says the surviving failure is a real case pressed into service for a proposition it does not carry, and the only thing that finds it is a person reading the case. The Institute’s Citation verification and Research workflow pages cover how firms build that step so it survives a deadline, and how often legal AI still gets it wrong takes up the error-rate question across tasks beyond research.

For informational purposes only. Not legal advice and not ethics advice. Professional conduct rules are adopted state by state and diverge, and this record changes monthly. Anything here that reads as a holding should be checked against your own jurisdiction before it is relied on.

Related

The practice area

AI adoption conciergeorientation · not legal or ethics advice
Happy to. Tell me roughly how big the firm is and what it already pays for — Microsoft 365, Google Workspace, a practice-management system — because the honest answer to most AI questions at a firm your size starts with what you have already bought rather than what you should go and buy.