What did the controlled study actually find?
That the effect depends almost entirely on the task, and that some tasks got worse. The UK Department for Business and Trade published its Microsoft 365 Copilot evaluation in August 2025: 1,000 licences, a diary study with 300 participants, 19 interviews, and — the part that distinguishes it — observed task sessions scored against a control group.
Most published Copilot evaluations rely on self-reported time savings. This one watched people do the work and scored the output, which is why its results diverge so sharply from the self-reported picture.
Its own summary conclusion was that the evaluators did not find robust evidence to suggest the time savings were leading to improved productivity.
Which tasks got worse with Copilot?
Data analysis in Excel and slide production, both substantially. On scored observed sessions, Excel data analysis came in at 1.5 out of 5 for accuracy with Copilot against 2.7 without, with quality showing the same 1.5 against 2.7 split, and the Copilot sessions took longer — about 25 minutes against about 20.
PowerPoint was faster and considerably worse: 1.5 against 5.0 on accuracy and 1.0 against 2.0 on quality, in roughly half the time. Faster production of materially weaker output is a specific and underappreciated failure mode, because the speed is visible to the person doing it and the quality gap is not.
In the diary study’s adjusted mean hours saved per task, image generation came out at minus 0.5 and scheduling at minus 0.6. Some tasks cost time.
Which tasks got dramatically better?
Summarising long documents, by a wide margin. On the observed sessions, report summarisation took about 12 minutes 37 seconds with Copilot against about 41 minutes 34 seconds without, at higher accuracy — 4.0 against 2.5 — and higher quality.
Drafting a report showed the largest adjusted mean time saving in the diary study at 1.3 hours, with research summarisation, meeting summarisation and information search clustered at 0.7 to 0.8. Email writing was 0.2, which is worth knowing given how often email drafting is the demonstration.
The pattern generalises well beyond this one product: reading and compressing is where these systems are strong, and judgment-laden editing and numerical reasoning are where they are weak.
Did the study say anything specific about legal work?
Yes, and it is not encouraging on the general case. Broken down by role type, the evaluation found that roles requiring high levels of contextual awareness, nuance and attention to detail — naming legal and policy roles specifically — found Microsoft 365 Copilot less suitable. One lawyer quoted in the study put it as: in a work context, especially dealing with legal text, it is important that every word is right.
That is a finding a firm should hold alongside the summarisation result rather than instead of it. The same tool that compresses a long report well is the one that has to be right about every word in a definition, and those are different jobs.
Who benefited most, and why does that matter?
Neurodiverse, dyslexic and non-native-English users, and they were the only group with statistically significant gains. Neurodiverse respondents were statistically more satisfied at 90% confidence and more likely to recommend the tool at 95%; any health condition or disability correlated with higher recommendation; non-native English speakers reported similar benefits. One dyslexic user described it as having made them more confident in reporting work.
That is the most defensible business case in the whole evaluation, and it is the one firms almost never lead with. It is a specific, measured, non-vendor claim about a real group of people, rather than a general productivity assertion the evidence does not support.
The training finding sits alongside it and is nearly as useful: self-led training produced significant satisfaction gains at two hours or more, while formal departmental training produced no significant effect below five hours.
Why do self-reported numbers look so much better than measured ones?
Because people report how the work felt, and these tools make work feel faster whether or not it is. HMRC’s Phase III deployment, reported 9 July 2026 across 3,500 licences, produced self-reported savings of two to three percent of the working week, roughly 60 minutes, with satisfaction at 7.1 out of 10 — and 46% of non-users citing security and data privacy as their main reason for not using it.
There is a measurable quality cost that self-reports miss entirely. BetterUp Labs with the Stanford Social Media Lab, published in Harvard Business Review on 22 September 2025, surveyed 1,150 US full-time workers and found 41% had received "workslop" — AI output that looks like work and is not — costing an average of 1 hour 56 minutes to fix per instance, roughly $186 per worker per month. That cost lands on the recipient, so it never appears in the sender’s self-report.
The Institute’s Adoption & Change area covers how to run a pilot that can actually produce a negative result.