Benchmark review · evidence before ranking

AI Benchmark Procurement Checklist for Legal Teams

A benchmark can help a legal team choose what to test. It cannot, by itself, establish matter accuracy, confidentiality, completed-task economics or safe deployment.

Direct answer

Use a benchmark in procurement only after recording its owner, dataset, task, model version, tools, prompt or agent setup, grader, date, cost and limitations. Then run the candidate on one approved legal workflow with the same sources, stop conditions and human-review standard. A leaderboard can shortlist a model; it cannot approve the workflow.

Six-field review frame for testing whether an AI benchmark can support a legal procurement decision
Review frame. A leaderboard becomes useful only after its dataset, version, tools and review conditions are recorded.

Start with a benchmark record, not the rank

FieldRecordProcurement question
Owner and datePublisher, version and checked dateWho can explain or update the result?
DatasetPublic, private or mixed; size and task familiesDoes it resemble the firm’s work?
CandidateExact model ID, provider route and availabilityCan the buyer purchase the tested configuration?
Run setupPrompt, tools, retrieval, agent scaffold and attempt countWhich surrounding system produced the score?
GraderAutomated rule, model grader or human rubricWhat does a pass actually mean?
EconomicsLatency, tokens, tool calls and failed runsWhat was omitted from the headline number?

Do not compare unlike benchmark names or scales

Two percentages can measure different things. The current Gemini legal material, for example, includes results attributed to different benchmark sources and setups. The existing Gemini benchmark provenance guide separates those records. Procurement should compare values only when the benchmark, dataset version, run configuration and scoring rule are materially compatible.

A higher score inside one disclosed chart can support a narrow statement about that chart. It does not establish that the model is more accurate on every practice area, language, jurisdiction, document type or agent workflow. Preserve the source’s own limitation language and mark missing setup fields Unknown.

Test dataset and contamination fit

Ask whether the tasks are public, private or recently refreshed; whether answer keys or close variants may have appeared in training data; and whether the benchmark reports contamination controls. A private dataset can reduce one exposure but also makes independent inspection harder. A public dataset improves reproducibility but may be familiar to models. Neither label resolves the issue alone.

Map the benchmark’s task families against the proposed workflow. Legal research, issue spotting, clause extraction, drafting, tool use and final-answer grading create different failure modes. If the dataset does not exercise the sources, permissions or abstentions that matter to the buyer, treat the result as discovery evidence rather than decision evidence.

Record the system around the model

Agent benchmarks often measure a model plus a scaffold: search, retrieval, browser or code tools, prompts, retry logic, context construction and graders. Record every component the publisher discloses. A firm may not be able to buy or reproduce the same route, and a managed legal product may add controls or failure modes not present in the benchmark.

Version pinning matters. A family name or latest alias can move after publication. Record the tested identifier, date and provider surface. If the benchmark author reruns the suite, preserve both snapshots rather than silently replacing the earlier procurement record.

Translate the score into one internal test

  1. Select one repeatable legal task and an approved source packet.
  2. Define expected material issues, required citations, prohibited actions and abstention cases.
  3. Run the current baseline and candidate with controlled tools and permissions.
  4. Have the same qualified reviewer score material corrections, unsupported assertions and completion time.
  5. Record run cost, retries, latency and reviewer minutes.
  6. Reject the candidate when a material stop condition fails, even if the public benchmark rank is higher.

Measure completed work, including human review

A model can improve a benchmark score while creating more review work. Procurement should measure the completed task: model cost, tool and retrieval cost, integration burden, failed runs, escalation rate and reviewer time. Keep the mean and the worst material failure visible when the sample is small.

Human review is not a generic disclaimer. Name the reviewer role, what they must verify and the evidence they can inspect. A final decision that cannot be reconstructed from sources, tool traces and corrections is weak procurement evidence regardless of response quality.

Close with a bounded procurement decision

The decision record should state what the public benchmark supports, what remains Unknown, which internal test was run, which data class and route were approved, who reviewed the result and when the record expires. Do not call the score a certification, a legal-accuracy rate or proof of compliance. NIST’s AI RMF is voluntary and can organize the evaluation; it does not convert a benchmark into approval.

FAQ

Can a legal AI benchmark identify the best model for a law firm?

It can identify candidates for a specific test. It cannot establish a universal winner without matching the firm’s task, sources, tools, permissions and review standard.

What benchmark fields matter most in procurement?

Record the owner, dataset, task, exact model version, system setup, grader, date, cost and limitations before using the result.

Is a vendor-reported benchmark independent evidence?

No. Label vendor-reported results as vendor evidence. Benchmark-author results have their own methodology and limitations and still need a firm-specific test.

Sources checked

Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.