There is real independent evidence here, which is rare in this market. It is worth reading properly rather than through a headline.
Start a conversation with the AI Adoption Concierge, already scoped to what the accuracy studies show. Pick a starting point, or describe your situation directly.
Most claims about AI accuracy in legal work come from the companies selling it. Legal research is the exception: independent researchers have tested the leading commercial tools and published what they found. The Stanford work reported hallucination rates between roughly 17% and 33% across the products examined — the best-performing system answering about 65% of queries accurately, another about 42% while hallucinating roughly twice as often — and specifically assessed vendor claims that retrieval architectures prevent hallucination, concluding those claims were overstated. Vendors have since published improved figures, including citation hallucination in low single digits for at least one product. Reading these numbers well means understanding what each measured, because they are not measuring the same thing.
Six questions that determine whether two numbers are comparable. Usually they are not.
Independent researchers or the vendor. Both can be informative; they are not equivalent.
A fabricated citation, a mischaracterised holding, or any unsupported statement. Definitions vary hugely and drive the number.
Hard questions, typical questions, or questions the tool was built for. This alone can move a result by a wide margin.
These products change between releases. A result from eighteen months ago describes a different system.
Against another tool, against a lawyer, or against nothing. Absolute rates without a baseline are hard to act on.
Citation existence, holding accuracy, or overall answer quality — three very different bars.
How firms use the evidence.
It is the only place in legal AI with genuinely independent measurement. Use it.
Both major vendors marketed retrieval as preventing hallucination. Independent testing found those claims overstated. Whatever the current rate for any product, "grounded" does not mean "verified" — and a firm that treats it that way has removed its own safety net.
Treat any figure as describing the system at the time it was tested, and attach the date whenever you quote one. These products change substantially between releases, and vendors have published improvements — citation hallucination reported in low single digits for at least one product, down from materially higher figures the year before. What has not changed is the structural point: the rate is not zero, no vendor claims it is zero, and a workflow that assumes zero is unsound regardless of which product is in use.
Mostly because they measure different things on different questions, not because anyone is lying. A vendor testing whether cited cases exist will report a much better number than a researcher testing whether the answer is legally correct. Question difficulty matters enormously — a set drawn from the tool's strongest area produces a different result from one designed to probe edges. Neither figure is meaningless; they answer different questions, and comparing them directly is what produces confusion.
If the tool will be used for substantive work, yes, and it is cheaper than it sounds. Twenty questions from your own practice, with answers you already know, scored by someone competent, will tell you more about the tool in your context than any published study — because published sets are general and your practice is not. It also gives the firm its own evidence, which is what you will want if anyone later asks how the decision to adopt was made.
It means you can calibrate it, not remove it. A meaningfully lower error rate might justify lighter review of internal working product while leaving full verification on anything filed or sent to a client — that is a sensible risk-tiered response. What it never justifies is dropping the check on court filings, because the consequence of the residual error is unchanged by its becoming rarer. Rare errors are also harder to catch, precisely because reviewers stop expecting them.
Describe what you are being told by vendors. The Institute will help you test the claims.