home  /  legal research with AI  /  what the accuracy studies show
research · ai for legal practice

What the accuracy studies show.

There is real independent evidence here, which is rare in this market. It is worth reading properly rather than through a headline.

begin here

Where is your firm?

Start a conversation with the AI Adoption Concierge, already scoped to what the accuracy studies show. Pick a starting point, or describe your situation directly.

AI Adoption Conciergewhat the accuracy studies show · orientation, not legal or ethics advice
I can tell you what the independent studies measured and how to read the numbers — I won't tell you which product to buy. Tell me what claims you are being shown and what research your firm does.

Most claims about AI accuracy in legal work come from the companies selling it. Legal research is the exception: independent researchers have tested the leading commercial tools and published what they found. The Stanford work reported hallucination rates between roughly 17% and 33% across the products examined — the best-performing system answering about 65% of queries accurately, another about 42% while hallucinating roughly twice as often — and specifically assessed vendor claims that retrieval architectures prevent hallucination, concluding those claims were overstated. Vendors have since published improved figures, including citation hallucination in low single digits for at least one product. Reading these numbers well means understanding what each measured, because they are not measuring the same thing.

mechanisms

How to read an accuracy claim.

Six questions that determine whether two numbers are comparable. Usually they are not.

Who ran it

Independent researchers or the vendor. Both can be informative; they are not equivalent.

What counted as a hallucination

A fabricated citation, a mischaracterised holding, or any unsupported statement. Definitions vary hugely and drive the number.

What the question set was

Hard questions, typical questions, or questions the tool was built for. This alone can move a result by a wide margin.

When it ran

These products change between releases. A result from eighteen months ago describes a different system.

What the comparison was

Against another tool, against a lawyer, or against nothing. Absolute rates without a baseline are hard to act on.

What was actually measured

Citation existence, holding accuracy, or overall answer quality — three very different bars.

methodology

What the evidence shows — and what we examine.

How firms use the evidence.

Calibrate the verification stepSet review depth to the observed error rate rather than to how confident the output reads.
Run your own question setPublished studies test general questions; your practice has specific ones. A small internal set is more informative.
Record what you foundYour own testing is the evidence you will want when explaining the firm's process later.
Re-test on major releasesCapability moves. A conclusion reached against last year's version is a historical fact, not a current one.
what's at stake

Why the evidence matters.

It is the only place in legal AI with genuinely independent measurement. Use it.

how much verification is warranted realistic expectations set internally what to tell partners and clients a defensible basis for the firm's process resistance to vendor claims honest productivity expectations

The finding that matters is about the architecture.

Both major vendors marketed retrieval as preventing hallucination. Independent testing found those claims overstated. Whatever the current rate for any product, "grounded" does not mean "verified" — and a firm that treats it that way has removed its own safety net.

common questions

The evidence — practical questions.

Are the published rates still current?

Treat any figure as describing the system at the time it was tested, and attach the date whenever you quote one. These products change substantially between releases, and vendors have published improvements — citation hallucination reported in low single digits for at least one product, down from materially higher figures the year before. What has not changed is the structural point: the rate is not zero, no vendor claims it is zero, and a workflow that assumes zero is unsound regardless of which product is in use.

Why do vendor numbers differ so much from independent ones?

Mostly because they measure different things on different questions, not because anyone is lying. A vendor testing whether cited cases exist will report a much better number than a researcher testing whether the answer is legally correct. Question difficulty matters enormously — a set drawn from the tool's strongest area produces a different result from one designed to probe edges. Neither figure is meaningless; they answer different questions, and comparing them directly is what produces confusion.

Should we run our own testing?

If the tool will be used for substantive work, yes, and it is cheaper than it sounds. Twenty questions from your own practice, with answers you already know, scored by someone competent, will tell you more about the tool in your context than any published study — because published sets are general and your practice is not. It also gives the firm its own evidence, which is what you will want if anyone later asks how the decision to adopt was made.

Does a low hallucination rate mean we can reduce verification?

It means you can calibrate it, not remove it. A meaningfully lower error rate might justify lighter review of internal working product while leaving full verification on anything filed or sent to a client — that is a sensible risk-tiered response. What it never justifies is dropping the check on court filings, because the consequence of the residual error is unchanged by its becoming rarer. Rare errors are also harder to catch, precisely because reviewers stop expecting them.

related

Related specialization areas & resources.

Read the evidence properly.

Describe what you are being told by vendors. The Institute will help you test the claims.

AI adoption conciergeorientation · not legal or ethics advice
I can tell you what the independent studies measured and how to read the numbers — I won't tell you which product to buy. Tell me what claims you are being shown and what research your firm does.