home  /  insights  /  does-copilot-make-lawyers-faster
Adoption & Change

Does Microsoft 365 Copilot actually make people faster?

The only UK government evaluation with a control group found Copilot users doing spreadsheet analysis were slower and materially less accurate. It also found one group with a clear statistically significant benefit, and that is the more useful finding.

September 4, 2026 · 5 min read

The short answer

On some tasks yes, dramatically, and on others it made people measurably worse. The UK Department for Business and Trade evaluation of August 2025 — 1,000 licences, a diary study of 300, 19 interviews, and observed task sessions scored against a control group — found Copilot users doing data analysis in Excel scored 1.5 out of 5 on accuracy against 2.7 without it, while taking longer, whereas report summarisation took roughly a third of the time at higher accuracy. It also found that roles requiring high contextual awareness and attention to detail, naming legal and policy roles specifically, found the tool less suitable. The group with statistically significant benefit was neurodiverse, dyslexic and non-native-English users, which is both the honest caution and the most defensible business case.

What this article establishes

  • The UK Department for Business and Trade evaluation (August 2025) is the only one of the major studies with observed task sessions against a control group.
  • Excel data analysis scored 1.5/5 on accuracy with Copilot against 2.7/5 without, and took longer; report summarisation took roughly a third of the time at higher accuracy.
  • The study named legal and policy roles as finding the tool less suitable, because of the contextual precision the work requires.
  • Neurodiverse, dyslexic and non-native-English users were the group with statistically significant satisfaction gains — the strongest evidence-backed business case in the study.

What did the controlled study actually find?

That the effect depends almost entirely on the task, and that some tasks got worse. The UK Department for Business and Trade published its Microsoft 365 Copilot evaluation in August 2025: 1,000 licences, a diary study with 300 participants, 19 interviews, and — the part that distinguishes it — observed task sessions scored against a control group.

Most published Copilot evaluations rely on self-reported time savings. This one watched people do the work and scored the output, which is why its results diverge so sharply from the self-reported picture.

Its own summary conclusion was that the evaluators did not find robust evidence to suggest the time savings were leading to improved productivity.

Which tasks got worse with Copilot?

Data analysis in Excel and slide production, both substantially. On scored observed sessions, Excel data analysis came in at 1.5 out of 5 for accuracy with Copilot against 2.7 without, with quality showing the same 1.5 against 2.7 split, and the Copilot sessions took longer — about 25 minutes against about 20.

PowerPoint was faster and considerably worse: 1.5 against 5.0 on accuracy and 1.0 against 2.0 on quality, in roughly half the time. Faster production of materially weaker output is a specific and underappreciated failure mode, because the speed is visible to the person doing it and the quality gap is not.

In the diary study’s adjusted mean hours saved per task, image generation came out at minus 0.5 and scheduling at minus 0.6. Some tasks cost time.

Which tasks got dramatically better?

Summarising long documents, by a wide margin. On the observed sessions, report summarisation took about 12 minutes 37 seconds with Copilot against about 41 minutes 34 seconds without, at higher accuracy — 4.0 against 2.5 — and higher quality.

Drafting a report showed the largest adjusted mean time saving in the diary study at 1.3 hours, with research summarisation, meeting summarisation and information search clustered at 0.7 to 0.8. Email writing was 0.2, which is worth knowing given how often email drafting is the demonstration.

The pattern generalises well beyond this one product: reading and compressing is where these systems are strong, and judgment-laden editing and numerical reasoning are where they are weak.

Yes, and it is not encouraging on the general case. Broken down by role type, the evaluation found that roles requiring high levels of contextual awareness, nuance and attention to detail — naming legal and policy roles specifically — found Microsoft 365 Copilot less suitable. One lawyer quoted in the study put it as: in a work context, especially dealing with legal text, it is important that every word is right.

That is a finding a firm should hold alongside the summarisation result rather than instead of it. The same tool that compresses a long report well is the one that has to be right about every word in a definition, and those are different jobs.

Who benefited most, and why does that matter?

Neurodiverse, dyslexic and non-native-English users, and they were the only group with statistically significant gains. Neurodiverse respondents were statistically more satisfied at 90% confidence and more likely to recommend the tool at 95%; any health condition or disability correlated with higher recommendation; non-native English speakers reported similar benefits. One dyslexic user described it as having made them more confident in reporting work.

That is the most defensible business case in the whole evaluation, and it is the one firms almost never lead with. It is a specific, measured, non-vendor claim about a real group of people, rather than a general productivity assertion the evidence does not support.

The training finding sits alongside it and is nearly as useful: self-led training produced significant satisfaction gains at two hours or more, while formal departmental training produced no significant effect below five hours.

Why do self-reported numbers look so much better than measured ones?

Because people report how the work felt, and these tools make work feel faster whether or not it is. HMRC’s Phase III deployment, reported 9 July 2026 across 3,500 licences, produced self-reported savings of two to three percent of the working week, roughly 60 minutes, with satisfaction at 7.1 out of 10 — and 46% of non-users citing security and data privacy as their main reason for not using it.

There is a measurable quality cost that self-reports miss entirely. BetterUp Labs with the Stanford Social Media Lab, published in Harvard Business Review on 22 September 2025, surveyed 1,150 US full-time workers and found 41% had received "workslop" — AI output that looks like work and is not — costing an average of 1 hour 56 minutes to fix per instance, roughly $186 per worker per month. That cost lands on the recipient, so it never appears in the sender’s self-report.

The Institute’s Adoption & Change area covers how to run a pilot that can actually produce a negative result.

For informational purposes only. Not legal advice and not ethics advice. Professional conduct rules are adopted state by state and diverge, and this record changes monthly. Anything here that reads as a holding should be checked against your own jurisdiction before it is relied on.

Related

The practice area

AI adoption conciergeorientation · not legal or ethics advice
Happy to. Tell me roughly how big the firm is and what it already pays for — Microsoft 365, Google Workspace, a practice-management system — because the honest answer to most AI questions at a firm your size starts with what you have already bought rather than what you should go and buy.