home  /  legal research with AI  /  grounding & retrieval
research · ai for legal practice

Grounding and retrieval.

Retrieval solves the most visible failure and leaves the more dangerous one untouched.

begin here

Where is your firm?

Start a conversation with the AI Adoption Concierge, already scoped to grounding & retrieval. Pick a starting point, or describe your situation directly.

AI Adoption Conciergegrounding & retrieval · orientation, not legal or ethics advice
Tell me which tools you are looking at and what research your firm actually does. I'll help you work out what to test — the useful tests are not the ones vendors demo.

A grounded system retrieves real documents and answers from them, rather than producing text from training alone. In legal research that distinction is significant: it largely removes the citation to a case that never existed, which is the failure that produces sanctions. What it does not remove is the harder problem — a real case retrieved and then characterised wrongly, a holding overstated, a synthesis that no retrieved document actually supports. Those errors arrive attached to genuine citations that check out, which makes them considerably harder for a reviewer to catch than a fabricated case would be. Understanding which failure a given architecture addresses is what lets a firm place its verification effort where it will find something.

mechanisms

What grounding does and does not fix.

The first two are largely addressed. The rest are not, and they are the ones that survive review.

Invented citations — largely fixed

If the answer is generated from retrieved documents, the documents exist. The most visible failure mode substantially recedes.

Fabricated quotations — reduced

Quoted text can be checked against the retrieved source, and better systems link directly to it.

Mischaracterised holdings — not fixed

A real case summarised as standing for something it does not. Survives retrieval entirely and passes a citation check.

Unsupported synthesis — not fixed

A conclusion drawn across several sources that none of them supports individually.

Retrieval gaps — not fixed

What the corpus does not contain cannot be retrieved, and the system rarely says so.

Currency — not fixed

Retrieved authority that has been reversed or superseded, unless the citator step is genuinely integrated.

methodology

What the evidence shows — and what we examine.

How to assess a system's grounding.

Click every citation throughTo the actual document in the vendor's corpus. If you cannot, the grounding claim is weaker than stated.
Establish corpus coverageWhich jurisdictions, which courts, how current, and whether secondary sources are included.
Ask for authority that does not existThe decisive test. A grounded system should find nothing and say so.
Check holdings against the sourceNot that the case exists — that it says what the summary claims. This is the surviving failure.
what's at stake

What the architecture determines.

Mainly where the firm should spend its verification effort, which is a finite resource.

where verification effort should go which errors will reach a reviewer how much time the tool really saves which jurisdictions it can serve what researchers must be taught to catch exposure on anything filed

The dangerous error passes the citation check.

A fabricated case is caught by anyone who looks it up. A real case cited for a holding it does not contain survives that check completely — and it is the error grounding does nothing to prevent.

common questions

Grounding — practical questions.

How do I tell whether a tool is genuinely grounded?

Follow the citations. In a genuinely grounded system every authority in the answer should link to a document the vendor actually holds, and you should be able to open it and read the passage the answer relied on. Where citations are presented as plain text with no link, or where the link goes to a search result rather than the document, the grounding claim is doing less work than it appears. Asking the vendor which corpus is retrieved from, and how it is kept current, usually settles it quickly.

Does a bigger corpus mean better answers?

Not straightforwardly, and coverage matters more than raw size. What determines whether a system can answer your question is whether it holds the jurisdictions, courts and time periods your practice needs, and how current it is — a vast corpus that is thin on your state's intermediate appellate decisions is not useful to you. Retrieval quality also matters independently: a system that holds the right document but does not surface it produces the same outcome as one that does not hold it.

What happens when the answer is not in the corpus?

This is the behaviour worth testing before buying anything, because it varies enormously. A well-built system reports that it found no supporting authority. A poorly built one generates a plausible answer from training rather than retrieval, and presents it in the same confident register as a grounded one. The user cannot distinguish the two from the output. Putting a question with no answer to the tool during evaluation tells you which kind you are dealing with in about a minute.

Is a general-purpose model with web search equivalent?

For legal research, generally not, and the difference is the corpus rather than the model. General web search reaches what is publicly indexed, which for case law is uneven, frequently secondary, and often stale — and it has no citator. Purpose-built legal tools retrieve from licensed primary law with subsequent history attached. That said, general models can be better at reasoning over material you supply, which is a different task from finding authority, and firms conflate the two more often than is helpful.

related

Related specialization areas & resources.

Understand what the architecture fixes.

Describe the tools you are evaluating. The Institute will help you test them properly.

AI adoption conciergeorientation · not legal or ethics advice
Tell me which tools you are looking at and what research your firm actually does. I'll help you work out what to test — the useful tests are not the ones vendors demo.