Licence counts and login rates make every deployment look successful. They are also the only numbers most firms have.
Start a conversation with the AI Adoption Concierge, already scoped to measuring value. Pick a starting point, or describe your situation directly.
Estimates of the time generative AI frees are large and widely quoted — figures on the order of 240 hours per lawyer per year appear in the analyses — but firm-level realisation is consistently lower than per-task studies imply, because verification, uneven adoption and partial use absorb much of it. That gap between theoretical and realised value is exactly what a firm needs to measure, and almost none can, because nothing was recorded before rollout. The instrumentation problem is more tractable than it looks: a small number of honest measures, captured on a baseline and tracked, will tell a firm more than a dashboard of usage statistics that only ever go up.
The useful measures are harder to collect, which is precisely why the useless ones dominate.
What proportion of output was usable without substantial redoing. The single most informative measure.
Time on the whole task including verification — not time to first draft, which flatters.
Scored by someone who did not produce it, on a simple scale, against the previous standard.
What proportion of eligible work uses it, rather than how many people logged in once.
Errors caught in review. Rising counts may mean the review is working, not that things are worse.
Crude, honest, and a better predictor of durable adoption than most instrumented metrics.
How to instrument it without a project.
Whether the next investment is argued from evidence or from enthusiasm.
Without a before, every after is an anecdote. Ten matters measured on rework rate and elapsed time in the fortnight before rollout is a small cost that makes every subsequent claim defensible — and it cannot be recovered afterwards.
Treat it as directional rather than as a number to plan on. Analyses putting the time freed on the order of 240 hours per lawyer per year exist and are widely cited, but they depend heavily on practice mix, on how completely the tools are adopted, and on what is counted. Firm-level realisation is reliably lower, because verification consumes some of it, adoption is uneven, and much of the theoretical saving sits in work the firm does not do much of. The distribution matters more than the total: the saving concentrates in specific matter types, and that is where to look.
Compare within a matter type rather than across the firm, and accept a rough comparison rather than chasing a rigorous one. No two matters are identical, but document review of a similar corpus, or first drafts of a familiar instrument, are comparable enough for a directional read — and directional is what the decision needs. Firms that hold out for methodological rigour typically end up measuring nothing, which is strictly worse than an imperfect number honestly labelled.
At minimum the firm needs to be able to see which work involved AI, or it cannot measure anything or answer a client question about it. Many firms have added a simple flag or task code. Whether that flows into what is billed is a separate question with fee-reasonableness implications, and it should be settled deliberately rather than as a by-product of a reporting change — the two decisions get conflated surprisingly often, and the reporting one is easier.
That is a valuable result and should be reported as one, because the alternative is paying for something indefinitely on the strength of an assumption. It also usually points somewhere specific: the tool is wrong for the task, the workflow is wrong, the training did not land, or the task was not actually a bottleneck. A firm that can distinguish those has learned something worth more than the licence. A firm that suppresses the finding will repeat it with the next product.
Describe what you are deploying and where. The Institute will help you decide what to measure.