home  /  library  /  large-legal-fictions
Independent studies

Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models

Matthew Dahl, Varun Magesh, Mirac Suzgun and Daniel E. Ho · 2024 · Journal of Legal Analysis; arXiv 2401.01301

Legal researchVerificationGovernanceFirms considering general-purpose AI tools for legal questions
Why it is on the shelf. Establishes the gap between general-purpose models and purpose-built legal tools, which is the evidence base for one of the most consequential policy decisions a firm makes: whether people may use consumer AI for work.

The Institute's reading

This is the study that should be read immediately before writing a firm AI policy, because it addresses the situation most firms are actually in rather than the one they plan for. Long before a firm licenses anything, its lawyers are using whatever is on their phone. The researchers profiled how general-purpose large language models behave on legal queries and found hallucination at rates far above anything reported for purpose-built legal tools — a majority of responses on many query types, rising higher on questions about lower courts and less prominent jurisdictions.

Two findings deserve particular attention from anyone drafting policy. The first is that error rates are not uniform: they climb sharply as questions move away from prominent federal authority toward state trial courts and specialised areas — which is to say, they are worst exactly where most practice happens. The second is that the models were poorly calibrated, tending to express confidence that did not track accuracy, and to accept the premise of a question containing a false one. A lawyer who asks about a case that does not exist may be told about it. That is not a failure a reader can catch by reading carefully, because nothing in the answer signals it.

Key propositions

  • General-purpose language models hallucinate on legal queries at rates far exceeding purpose-built legal research tools.
  • Error rates rise substantially for lower courts, less prominent jurisdictions and specialised areas.
  • The models are poorly calibrated: expressed confidence does not track accuracy.
  • They frequently accept and build on false premises embedded in a question.
  • The gap between general and legal-specific tools is large enough to be a policy-relevant distinction.

In practice

  • A firm policy that does not distinguish general-purpose tools from legal-specific ones is not addressing the main risk.
  • The practice areas most exposed are those furthest from prominent federal authority — most of them.
  • Training must cover that confident phrasing carries no information about correctness, because readers assume it does.

Where authorities disagree

General-purpose models have improved substantially since this testing and the specific rates should be read as historical. The contested question is how much of the gap remains: some argue that frontier models with retrieval now approach specialist tools, while others hold that the absence of a licensed primary-law corpus and a citator is a structural limit no amount of model capability closes. The Institute's position is that the distinction still matters for policy, and that a firm relying on it should test rather than assume.

AI adoption conciergeorientation · not legal or ethics advice
Happy to dig into it. What would you like to pressure-test from Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models: one of its propositions, how it applies to your situation, or where it disagrees with the rest of the shelf?