Responsibility for the output has not moved. The ability to inspect how it was produced has.
Start a conversation with the AI Adoption Concierge, already scoped to agentic systems & autonomy. Pick a starting point, or describe your situation directly.
An agentic system does not answer a question — it pursues a goal. It plans steps, runs searches, reads what it finds, drafts, checks its own work and calls other tools, looping until it decides it is done. For legal work the appeal is obvious, because a great deal of practice is exactly that kind of multi-step procedure. The difficulty is equally clear and it is not a technical one: professional responsibility for the output is unchanged, while the number of intermediate decisions a supervising lawyer would have to inspect to genuinely verify it has grown by an order of magnitude. The Law Society has named this directly as a widening gap between responsibility and auditability. It is not a transitional problem that better tools will quietly close.
Supervision is straightforward at the top and genuinely hard at the bottom.
One prompt, one answer. The lawyer sees exactly what was produced and reviews it. Conventional.
The system runs defined steps in order. Predictable, and each step is inspectable.
The system chooses its own steps toward a stated goal. The plan itself becomes something to review.
It calls search, documents and other systems. What it actually consulted is now a question with a non-obvious answer.
It reviews and revises its own output. Useful, and it means the reviewed version may hide the discarded reasoning.
It finishes without a checkpoint. This is where responsibility and auditability come apart.
How firms keep agentic work supervisable.
Not the duty. Only the difficulty of discharging it.
That is the gap, stated plainly. It is why the systems worth deploying are the ones that escalate decisions to a human and keep a readable record — not the ones that impress by needing nobody.
Yes, and mostly in narrow, bounded applications rather than the general autonomy the term suggests. Several large firms have deployed agentic tooling internally for things like multi-step research, document processing pipelines and diligence review, where the steps are well-defined and the output lands in front of a lawyer. The gap between that and a system running a matter is very wide. Treat "we use agentic AI" as a statement about a workflow, not a description of autonomous practice, and ask what specifically it does unattended.
By designing checkpoints in rather than hoping review catches things at the end. The workable pattern is to identify the decisions that actually matter in a given workflow and require human approval at each, so the system does the volume and the lawyer makes the judgements. This is less efficient than full autonomy and it is the version that fits the professional obligations as they stand. A supervisor reviewing only a final output from a long autonomous chain is performing review in name, because the errors that matter are upstream of what they can see.
More than feels necessary, because the value only appears when something has gone wrong. At minimum: the instruction given, the steps the system chose, what sources it consulted, what it produced at each stage, and who approved what. The reason is that if the output is challenged — by a client, a regulator or a court — the firm needs to reconstruct how it was produced, and a log containing only the final answer cannot do that. Confirm the retention period is long enough to outlive the limitation period on the underlying matter.
It is a question to put to your carrier rather than one to answer from first principles, and it is worth putting now rather than after an incident. The standard is still what a reasonably competent lawyer would do, and that standard adapts to available technology in both directions — over time it may become unreasonable not to use certain tools, just as it is unreasonable to rely on them uncritically now. What is new is the evidentiary problem: demonstrating adequate supervision of a system whose intermediate steps were not recorded is very difficult, which is a practical argument for logging.
Describe what you want it to do unattended. The Institute will help you place the checkpoints.