home  /  tool selection & evaluation  /  evaluation criteria
selection · ai for legal practice

Evaluation criteria.

Most procurement questions in this market are answered by the demo. The useful ones are answered by the contract and by a test on your own documents.

begin here

Where is your firm?

Start a conversation with the AI Adoption Concierge, already scoped to evaluation criteria. Pick a starting point, or describe your situation directly.

AI Adoption Conciergeevaluation criteria · orientation, not legal or ethics advice
Tell me what task you want the tool to improve and what you are considering. I'll help you build an evaluation — including the questions vendors do not volunteer. I won't rank named products.

Evaluating an AI tool is unlike evaluating other legal software, because the thing you are assessing is the quality of generated output rather than the presence of a feature. Feature checklists do not separate these products; behaviour under your own material does. Two questions carry most of the weight and neither is usually raised in a demo: what does the system do when it does not know, and what do the terms permit the vendor to do with what you put in. A tool that fabricates confidently and a tool that says it found nothing are indistinguishable on a feature list and completely different in practice.

mechanisms

What actually separates tools.

Roughly in order of how much they matter and how rarely they are asked about.

Grounding

Does output cite from a real retrieved corpus, or generate references? The single largest determinant of citation risk.

Behaviour at the limit

What it does when the answer is not there — declines, or produces something anyway with equal confidence.

Data handling terms

Whether inputs may be retained, used for training, or seen by subprocessors. Frequently eliminates candidates outright.

Integration

Whether it works inside the document system and practice management already in use, or demands a new tab and new habits.

Permissioning

Whether access respects matter-level ethical walls, which many general tools simply do not model.

Switching cost

What leaving costs in a market where capability leadership may change inside the contract term.

methodology

What the evidence shows — and what we examine.

How to test rather than watch a demo.

Test on your own mattersReal documents and real questions from your practice, not the vendor's curated examples.
Include questions with no answerThe most informative test in the set — it reveals whether the tool fabricates rather than declines.
Score against a defined markAgreed before the test, so the evaluation can fail rather than being rationalised into a purchase.
Read the terms before the demoData handling frequently disqualifies a product regardless of how well it performs.
what's at stake

What the choice determines.

Selection sets the risk profile everything downstream has to manage.

how much verification is needed whether client material can be used at all whether anyone actually adopts it licence cost against realised value months lost to a tool that goes unused how expensive it is to change course

Ask it something with no answer.

Put a question to the tool that has no supporting authority, and watch. A system that declines is telling you something a feature list never will. A system that invents an answer has told you exactly how much review it will require forever.

common questions

Evaluation — practical questions.

What does "grounded" actually mean?

That the system retrieves real documents and generates its answer from them, rather than producing text from its training alone. In practice the distinction shows up in whether a citation can be clicked through to a source the vendor actually holds. It reduces fabricated authority substantially — it does not eliminate the risk that a real retrieved case is characterised wrongly, which is why the proposition check survives regardless of the tool. Grounding is the difference between "this citation exists" and "this citation supports this point," and only the first is solved by architecture.

Which contract terms matter most?

Whether inputs may be used to train models, how long data is retained, who the subprocessors are, where the data sits, and what happens on termination. Professional obligations around client confidentiality bear directly on this — guidance including ABA Formal Opinion 512 addresses confidentiality in the context of generative AI, and a common reading is that client information should not go into tools whose terms permit training on it. Firms should have counsel read the terms rather than rely on marketing assurances about privacy, which are often about a different tier of the product.

How do we compare accuracy between tools?

On a fixed set of your own questions, scored by someone who knows the right answers, with the tools run blind where possible. Vendor benchmarks are not comparable across products and are rarely constructed on work resembling yours. The set should include questions where the answer is settled, questions where authority is split, and questions with no supporting authority at all. The third category separates products more sharply than the first two, and almost nobody includes it.

Should we wait, given how fast this is moving?

Waiting is itself a position with costs, and the market is unlikely to settle in a way that rewards it. The more defensible version of the instinct is to avoid long commitments rather than to avoid adoption — shorter terms, exportable data, capability categories rather than single-vendor dependence. Firms that have deferred entirely tend to find their people adopted consumer tools independently, which produces the confidentiality exposure without any of the governance.

related

Related specialization areas & resources.

Ask the questions the demo skips.

Describe what you are evaluating. The Institute will help you build the test.

AI adoption conciergeorientation · not legal or ethics advice
Tell me what task you want the tool to improve and what you are considering. I'll help you build an evaluation — including the questions vendors do not volunteer. I won't rank named products.