This is the study that should be read immediately before writing a firm AI policy, because it addresses the situation most firms are actually in rather than the one they plan for. Long before a firm licenses anything, its lawyers are using whatever is on their phone. The researchers profiled how general-purpose large language models behave on legal queries and found hallucination at rates far above anything reported for purpose-built legal tools — a majority of responses on many query types, rising higher on questions about lower courts and less prominent jurisdictions.
Two findings deserve particular attention from anyone drafting policy. The first is that error rates are not uniform: they climb sharply as questions move away from prominent federal authority toward state trial courts and specialised areas — which is to say, they are worst exactly where most practice happens. The second is that the models were poorly calibrated, tending to express confidence that did not track accuracy, and to accept the premise of a question containing a false one. A lawyer who asks about a case that does not exist may be told about it. That is not a failure a reader can catch by reading carefully, because nothing in the answer signals it.
General-purpose models have improved substantially since this testing and the specific rates should be read as historical. The contested question is how much of the gap remains: some argue that frontier models with retrieval now approach specialist tools, while others hold that the absence of a licensed primary-law corpus and a citator is a structural limit no amount of model capability closes. The Institute's position is that the distinction still matters for policy, and that a firm relying on it should test rather than assume.