home  /  tool selection & evaluation  /  pilots & proof of value
selection · ai for legal practice

Pilots and proof of value.

Most pilots succeed. That is the problem — a trial that cannot produce a negative result has not tested anything.

begin here

Where is your firm?

Start a conversation with the AI Adoption Concierge, already scoped to pilots & proof of value. Pick a starting point, or describe your situation directly.

AI Adoption Conciergepilots & proof of value · orientation, not legal or ethics advice
Tell me what task you would pilot on and who would take part. I'll help you set a measure and a pass mark — the two things most firm pilots skip.

The standard firm pilot recruits volunteers who are already interested, runs on convenient work, has no agreed measure of success, and concludes with a positive impression. It would have returned that result for almost any product. A pilot worth running has a defined task, a defined comparison, participants who are not all enthusiasts, and a pass mark set before it starts. It should also measure the thing the firm actually cares about — which is rarely raw speed and usually whether the work was good enough to use without redoing it.

mechanisms

What makes a pilot informative.

Each of these is routinely omitted, and each omission biases the result toward yes.

One defined task

A specific job, not "try it and see". Broad pilots produce broad impressions and no decision.

A pass mark set in advance

Agreed before the trial, so the result can be no rather than reinterpreted afterwards.

Participants who are not all keen

Include sceptics and ordinary users, or you have measured enthusiasm rather than the tool.

A comparison

Against the current way of doing it — otherwise there is no baseline and no claim to make.

Quality, not just speed

Output good enough to use without redoing. Time saved that produces unusable work is not saved.

Written up honestly

Including what did not work, which is the part that makes the next pilot better.

methodology

What the evidence shows — and what we examine.

How a pilot is run.

Pick a task with a clear edgeRepetitive, well-defined work with an obvious quality test beats an ambitious one nobody can score.
Define the measureTime, rework rate, and reviewer-rated quality — chosen before, not selected afterwards to fit the outcome.
Choose a mixed cohortSceptics included. Their objections are the most useful output of the exercise.
Report against the markA short written result that says pass or fail, and what would change the answer.
what's at stake

What a bad pilot costs.

Usually not the licence fee. Usually the year that follows it.

licences for a tool nobody opens credibility for the next initiative a year before anyone revisits it sceptics confirmed in their scepticism no baseline for future comparison switching cost incurred for nothing

Recruit at least one sceptic.

A pilot staffed entirely by volunteers measures the volunteers. The colleague who thinks this is overhyped will find the failure modes fastest, and if the tool convinces them it will convince the firm.

common questions

Pilots — practical questions.

What should a pilot actually measure?

Rework rate is the most informative single measure and the least often used: what proportion of the output was good enough to use without substantial redoing. Time saved is worth capturing but misleads on its own, because a tool that halves drafting time and doubles review time has saved nothing. Reviewer-rated quality on a simple scale, applied by someone who did not produce the work, adds the dimension raw speed misses. Adoption during the pilot is also telling — people quietly stopping is a result.

Which task makes a good first pilot?

Something repetitive, well-bounded, internally facing, and easy to score. Summarising documents into a defined format, first-pass review against fixed criteria, or drafting routine correspondence all work. The temptation is to pilot on the most valuable and most complex work, because that is where the upside looks largest — but complex work is hard to score, high-risk to get wrong, and produces an ambiguous result that settles nothing.

How do we stop the pilot becoming permanent by default?

Set an end date and a decision point at the start, and name who decides. A surprising number of firm pilots simply continue: the trial licences roll over, usage drifts, and no one ever decides whether it worked. That is the worst outcome, because the firm bears the cost and the governance exposure without ever having concluded the tool is worth it. A written result at a fixed date — pass, fail, or extend for a stated reason — avoids it.

Should clients know their matters are in a pilot?

This is a question for the firm and its own advisers, and worth resolving before the pilot rather than during it. It touches client confidentiality, outside counsel guidelines that increasingly address AI use, and the communication duties discussed in guidance such as ABA Formal Opinion 512. Many firms structure early pilots on internal or non-confidential material specifically to avoid the question until the governance is settled, which is usually the simpler sequence.

related

Related specialization areas & resources.

Design a pilot that can fail.

Describe the task and who would take part. The Institute will help you set the measure.

AI adoption conciergeorientation · not legal or ethics advice
Tell me what task you would pilot on and who would take part. I'll help you set a measure and a pass mark — the two things most firm pilots skip.