What do the hallucination studies actually measure?
Two distinct things, and conflating them produces most of the confusion in this area. "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models" (Dahl, Magesun, Suzgun and Ho, Journal of Legal Analysis, 2024) tested general-purpose chatbots against more than 800,000 verifiable legal questions about federal cases and reported hallucination rates of 58% for GPT-4, 69% for GPT-3.5 and 88% for Llama 2.
Those are figures for general-purpose models with no retrieval layer, which is not what most firms are buying. The same study documented a behaviour that matters independently of the rate: models frequently failed to correct a user’s false legal premise and instead built on it.
How do the legal-specific research products compare?
Better, and not as much better as marketed. "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools" (Magesh and others) was the first preregistered empirical evaluation of proprietary retrieval-augmented legal research tools, appearing as a preprint in May 2024 and in the Journal of Empirical Legal Studies in 2025.
It reported approximately 17% for Lexis+ AI and approximately 33–34% for Westlaw AI-Assisted Research, against 43% for GPT-4 with no retrieval on the same queries. A third product hallucinated less but gave incomplete answers on more than 60% of queries, which is a different failure with similar consequences.
These are named products and named measured results rather than a ranking. The Institute does not rank tools, and a benchmark from one period is not a claim about how any product performs later.
Have the tools got better since those studies?
Measurably, to the point of beating an unaided lawyer on research accuracy. Vals AI’s October 2025 legal research benchmark found legal-specific AI products reaching 78–81% accuracy against a human lawyer baseline of 69%, with a general assistant using web search at 80%.
The synthesis position from the AI Law Librarians review of February 2026 is the one to carry: even the best legal AI tools still err roughly 15 to 25 percent of the time on research tasks, and degradation is worst in non-major jurisdictions and on complex multi-jurisdictional questions. The honest framing for attorneys is that these tools are better than an average unaided associate on speed and often on accuracy, and never citable without human verification.
What is misgrounding, and why is it worse than a fake case?
Misgrounding is a real, correctly cited authority offered for a proposition it does not actually support. It is worse than a fabricated case for a simple mechanical reason: the standard verification step catches fabrication and does not catch this.
An associate checking citations confirms that each case exists, is good law, and is cited accurately as to court and date. Every one of those checks passes on a misgrounded citation. Catching it requires reading the authority for what it holds, which is the expensive step people skip precisely because the citation looks clean.
The Stanford work identified this as a second failure class distinct from fabrication, and almost no practitioner-facing guidance addresses it. It is the single most under-covered risk in legal AI.
How large is the sanctions record now?
Damien Charlotin of HEC Paris maintains the AI Hallucination Cases database, which recorded 2,009 decisions as of 3 September 2026 — 1,378 of them in the United States, comprising 1,156 involving self-represented parties, 799 involving lawyers, 31 involving judges and 15 involving experts. By nature of the defect: 1,667 fabricated authorities, 838 mischaracterised, 544 false quotes.
The trajectory is the part to notice. Roughly 200 in mid-2025, 719 in January 2026, 1,227 in early April 2026, 1,598 on 9 June 2026. The curve is accelerating rather than flattening, three years after the first widely reported sanctions in Mata v. Avianca (S.D.N.Y., June 2023).
The instructive recent example is Wadsworth v. Walmart (D. Wyo., 24 February 2025), where attorneys filed motions in limine citing nine cases of which eight did not exist, generated by an internal AI tool. The drafter was fined $3,000 and had pro hac vice admission revoked; two signing attorneys were fined $1,000 each. The signing attorneys had not done the drafting.
What does a defensible verification step actually look like?
It reads the authority rather than confirming the citation, and it is assigned to a named person rather than to a process. The distinguishing feature of a defensible checkpoint is that it can produce a negative result: if the reviewer has no realistic way to conclude "this does not support the proposition", the step is documentation rather than verification.
Two practical corollaries follow from the sanctions record. The signing attorney carries the exposure regardless of who drafted, which Wadsworth v. Walmart made concrete. And an internal tool offers no protection whatever — the fabrications in that matter came from one.
The Institute’s Verification & Quality Control area covers workflow design, and Legal Research with AI covers where retrieval helps and where it does not.