What does the evidence say about firms building their own AI?
The most cited finding is that internal builds fail more often than purchases. MIT NANDA’s "The GenAI Divide: State of AI in Business 2025" (August 2025) — 150 leader interviews, a 350-employee survey and 300 public deployments — reported that internally built systems succeeded roughly one-third as often as buying from specialised vendors, with buying showing a 67% success rate.
That build-versus-buy differential is the most useful number in the study, and it is not the number that got the coverage.
Is it true that 95% of AI pilots fail?
That figure is widely over-read and should be cited carefully. The MIT NANDA study reported that about 95% of enterprise generative-AI pilots stalled with no measurable profit-and-loss impact, which is a statement about pilots reaching demonstrated P&L impact — not a statement that the technology did not work in 95% of cases.
It is also a working-paper claim rather than settled fact, and it deserves that label whenever it is quoted. The honest version is less dramatic and more useful: most pilots do not produce a measurable financial result, largely because most pilots are not designed to be able to.
Which points at a design problem worth naming. A pilot with no baseline, no control and no pre-committed success measure cannot produce a negative result. It can only produce enthusiasm or silence.
Why do internal builds fail more often than purchases?
Because the hard parts are not the model. They are data hygiene, permissions, evaluation and ownership — the unglamorous infrastructure that a vendor has already paid for and that a firm building in-house has to fund out of its own overhead, indefinitely.
Ownership is the one firms underestimate most. A bought system has a vendor whose business depends on it continuing to work. A built system has whoever built it, until that person is busy on a matter, or leaves.
What does "configure" mean in practice?
It means switching on and tuning capability the firm has already bought rather than acquiring or constructing new capability. For most firms the largest untapped surface is inside an existing Microsoft 365 or Google Workspace agreement, followed by whatever the practice-management and document-management systems already ship.
This matters commercially as well as technically, because of a pricing pattern the Institute has documented across categories: the product tier that satisfies a firm’s confidentiality obligations is frequently priced for organisations ten times the buyer’s size. A three-lawyer firm can often afford a tool and not the tier that makes it defensible. Using what is already inside an enterprise agreement sidesteps that problem entirely, and costs nothing further.
When is building actually the right answer?
When the workflow is genuinely proprietary, the data to support it is already clean, and the firm can staff its maintenance for years rather than months. Those three conditions have to hold together; two out of three produces a system that works impressively for a quarter and then decays.
There is a middle path that is frequently the right one and rarely considered: connecting systems that already exist rather than building a new one. Practice-management, document-management and research systems increasingly ship permission-bound connectors, and wiring those together is a configuration exercise with a much shorter failure tail than construction.
One caution on that path. Where a connector is community-maintained rather than vendor-published, it may hold credentials to an entire practice-management system, and that is a serious trust decision rather than an installation step. Read the code, or do not run it.
Where do these projects actually die?
On data readiness, almost always, and earlier than anyone plans for. The MIT NANDA finding is substantially a data-readiness story dressed up as a technology story.
A law firm’s document estate is unusually hostile to retrieval: twenty years of document management with near-duplicates in the dozens, files called Document1.doc, fourteen versions of the same brief with no indication which was filed, PDFs of scans of faxes, matter numbers in four incompatible conventions, and an email archive nobody has indexed. No model fixes that. The assessment is boring, it is skipped almost universally, and it determines the outcome.
The Institute’s Automation & Building area covers the build-buy-configure decision in more depth.