Start with a benchmark record, not the rank
| Field | Record | Procurement question |
|---|---|---|
| Owner and date | Publisher, version and checked date | Who can explain or update the result? |
| Dataset | Public, private or mixed; size and task families | Does it resemble the firm’s work? |
| Candidate | Exact model ID, provider route and availability | Can the buyer purchase the tested configuration? |
| Run setup | Prompt, tools, retrieval, agent scaffold and attempt count | Which surrounding system produced the score? |
| Grader | Automated rule, model grader or human rubric | What does a pass actually mean? |
| Economics | Latency, tokens, tool calls and failed runs | What was omitted from the headline number? |
Do not compare unlike benchmark names or scales
Two percentages can measure different things. The current Gemini legal material, for example, includes results attributed to different benchmark sources and setups. The existing Gemini benchmark provenance guide separates those records. Procurement should compare values only when the benchmark, dataset version, run configuration and scoring rule are materially compatible.
A higher score inside one disclosed chart can support a narrow statement about that chart. It does not establish that the model is more accurate on every practice area, language, jurisdiction, document type or agent workflow. Preserve the source’s own limitation language and mark missing setup fields Unknown.
Test dataset and contamination fit
Ask whether the tasks are public, private or recently refreshed; whether answer keys or close variants may have appeared in training data; and whether the benchmark reports contamination controls. A private dataset can reduce one exposure but also makes independent inspection harder. A public dataset improves reproducibility but may be familiar to models. Neither label resolves the issue alone.
Map the benchmark’s task families against the proposed workflow. Legal research, issue spotting, clause extraction, drafting, tool use and final-answer grading create different failure modes. If the dataset does not exercise the sources, permissions or abstentions that matter to the buyer, treat the result as discovery evidence rather than decision evidence.
Record the system around the model
Agent benchmarks often measure a model plus a scaffold: search, retrieval, browser or code tools, prompts, retry logic, context construction and graders. Record every component the publisher discloses. A firm may not be able to buy or reproduce the same route, and a managed legal product may add controls or failure modes not present in the benchmark.
Version pinning matters. A family name or latest alias can move after publication. Record the tested identifier, date and provider surface. If the benchmark author reruns the suite, preserve both snapshots rather than silently replacing the earlier procurement record.
Translate the score into one internal test
- Select one repeatable legal task and an approved source packet.
- Define expected material issues, required citations, prohibited actions and abstention cases.
- Run the current baseline and candidate with controlled tools and permissions.
- Have the same qualified reviewer score material corrections, unsupported assertions and completion time.
- Record run cost, retries, latency and reviewer minutes.
- Reject the candidate when a material stop condition fails, even if the public benchmark rank is higher.
Measure completed work, including human review
A model can improve a benchmark score while creating more review work. Procurement should measure the completed task: model cost, tool and retrieval cost, integration burden, failed runs, escalation rate and reviewer time. Keep the mean and the worst material failure visible when the sample is small.
Human review is not a generic disclaimer. Name the reviewer role, what they must verify and the evidence they can inspect. A final decision that cannot be reconstructed from sources, tool traces and corrections is weak procurement evidence regardless of response quality.
Close with a bounded procurement decision
The decision record should state what the public benchmark supports, what remains Unknown, which internal test was run, which data class and route were approved, who reviewed the result and when the record expires. Do not call the score a certification, a legal-accuracy rate or proof of compliance. NIST’s AI RMF is voluntary and can organize the evaluation; it does not convert a benchmark into approval.
FAQ
Can a legal AI benchmark identify the best model for a law firm?
It can identify candidates for a specific test. It cannot establish a universal winner without matching the firm’s task, sources, tools, permissions and review standard.
What benchmark fields matter most in procurement?
Record the owner, dataset, task, exact model version, system setup, grader, date, cost and limitations before using the result.
Is a vendor-reported benchmark independent evidence?
No. Label vendor-reported results as vendor evidence. Benchmark-author results have their own methodology and limitations and still need a firm-specific test.
Sources checked
- Harvey LAB-AA leaderboard and method, checked 2026-09-03.
- Intelligence benchmarking methodology, checked 2026-09-03.
- Benchmark catalog and task descriptions, checked 2026-09-03.
- Gemini 3.8 Flash model card, checked 2026-09-03.
- AI Risk Management Framework, checked 2026-09-03.
Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.