Most procurement questions in this market are answered by the demo. The useful ones are answered by the contract and by a test on your own documents.
Start a conversation with the AI Adoption Concierge, already scoped to evaluation criteria. Pick a starting point, or describe your situation directly.
Evaluating an AI tool is unlike evaluating other legal software, because the thing you are assessing is the quality of generated output rather than the presence of a feature. Feature checklists do not separate these products; behaviour under your own material does. Two questions carry most of the weight and neither is usually raised in a demo: what does the system do when it does not know, and what do the terms permit the vendor to do with what you put in. A tool that fabricates confidently and a tool that says it found nothing are indistinguishable on a feature list and completely different in practice.
Roughly in order of how much they matter and how rarely they are asked about.
Does output cite from a real retrieved corpus, or generate references? The single largest determinant of citation risk.
What it does when the answer is not there — declines, or produces something anyway with equal confidence.
Whether inputs may be retained, used for training, or seen by subprocessors. Frequently eliminates candidates outright.
Whether it works inside the document system and practice management already in use, or demands a new tab and new habits.
Whether access respects matter-level ethical walls, which many general tools simply do not model.
What leaving costs in a market where capability leadership may change inside the contract term.
How to test rather than watch a demo.
Selection sets the risk profile everything downstream has to manage.
Put a question to the tool that has no supporting authority, and watch. A system that declines is telling you something a feature list never will. A system that invents an answer has told you exactly how much review it will require forever.
That the system retrieves real documents and generates its answer from them, rather than producing text from its training alone. In practice the distinction shows up in whether a citation can be clicked through to a source the vendor actually holds. It reduces fabricated authority substantially — it does not eliminate the risk that a real retrieved case is characterised wrongly, which is why the proposition check survives regardless of the tool. Grounding is the difference between "this citation exists" and "this citation supports this point," and only the first is solved by architecture.
Whether inputs may be used to train models, how long data is retained, who the subprocessors are, where the data sits, and what happens on termination. Professional obligations around client confidentiality bear directly on this — guidance including ABA Formal Opinion 512 addresses confidentiality in the context of generative AI, and a common reading is that client information should not go into tools whose terms permit training on it. Firms should have counsel read the terms rather than rely on marketing assurances about privacy, which are often about a different tier of the product.
On a fixed set of your own questions, scored by someone who knows the right answers, with the tools run blind where possible. Vendor benchmarks are not comparable across products and are rarely constructed on work resembling yours. The set should include questions where the answer is settled, questions where authority is split, and questions with no supporting authority at all. The third category separates products more sharply than the first two, and almost nobody includes it.
Waiting is itself a position with costs, and the market is unlikely to settle in a way that rewards it. The more defensible version of the instinct is to avoid long commitments rather than to avoid adoption — shorter terms, exportable data, capability categories rather than single-vendor dependence. Firms that have deferred entirely tend to find their people adopted consumer tools independently, which produces the confidentiality exposure without any of the governance.
Describe what you are evaluating. The Institute will help you build the test.