Downloadable utility · evidence before expansion

Legal AI Pilot Scorecard for Frontier Models

A pilot score should make a decision easier to audit. It should not compress a material legal failure into a reassuring average.

Direct answer

Freeze the baseline, define blocking failures, score every representative task and measure the completed result: material quality, source accuracy, reviewer corrections and minutes, model and tool cost, latency, permissions, stops, multilingual fit and rollback. Use the weighted score only after every blocking guardrail passes. A passing score is a dated pilot decision, not certification.

Legal AI pilot scorecard gates for baseline, quality, review effort, economics, permissions and rollback
Evidence gate. A score organizes a bounded decision; it does not certify a model or workflow.

Download the scorecard

Download the CSV template

The template includes task identity, baseline and candidate versions, criteria counts, blocking failures, source accuracy, material omissions, reviewer corrections and minutes, run cost, latency, refusal and stop results, multilingual notes, rollback result, decision and reviewer. No sample row is labeled as evidence.

Six gates before a weighted score

GateMinimum recordBlocking example
Frozen baselineProduct, account, model ID, prompt, sources, tools and dateThe comparator changes mid-pilot.
Material qualityExpected issues, citations, omissions and prohibited assertionsUnsupported authority or missed material protection.
Reviewer effortQualified reviewer, corrections and minutesNo accountable final reviewer.
Completed-task costModel, tools, retries and reviewer timeGenerated draft counted as completed work.
Permission and stop testAllowed tools, destinations, denials and abstentionsUnapproved read, write or send succeeds.
Rollback evidencePrior route, trigger, owner and tested recoveryThe workflow cannot return to the approved baseline.

Use explicit formulas

  • Criterion pass rate = passed criteria ÷ evaluated criteria × 100.
  • Strict task pass = 1 only when all blocking criteria and the defined minimum pass.
  • Source accuracy = verified cited propositions ÷ checked cited propositions × 100.
  • Completed-task cost = model + tools + retries + reviewer minutes ÷ 60 × loaded reviewer rate.
  • Review delta = candidate reviewer minutes − baseline reviewer minutes.

Do not average away a blocking failure. A high criterion rate with an invented authority still fails the task.

Choose a small representative set

Include ordinary work, long material, conflicting sources, missing information, adversarial instructions, tool failure, required abstention and one edge case from the actual practice. For multilingual work, include the languages and jurisdictions the firm will use, not a generic translation prompt.

Keep the same source packet, instructions, output format, tools and reviewer standard across the baseline and candidate. If a condition changes, start a new comparison row.

Decision states

StateMeaning
Approve bounded pilotAll blocking gates pass for the named workflow, data class, account and period.
RestrictThe route is useful only with narrower data, tools, users or review.
Revise and retestA remediable control or configuration failed.
RejectA material quality, permission, economics or rollback boundary fails.
UnknownRequired evidence was unavailable; absence of evidence is not a pass.

What the score cannot establish

The scorecard supports a workflow decision. It does not certify the model, preserve privilege, prove compliance, predict all matters or replace legal judgment. ABA guidance and NIST can inform governance, but the controlling professional obligations and risk assessment depend on the facts and jurisdiction.

FAQ

What should a legal AI pilot measure?

Measure material quality, source accuracy, reviewer corrections and time, completed-task cost, latency, permissions, stops, multilingual fit and rollback against a frozen baseline.

Can a weighted score offset a serious legal failure?

No. Define blocking criteria first. Unsupported authority, prohibited disclosure or unauthorized action should fail the task regardless of the average.

Does a passing score certify the AI model?

No. It supports a dated, bounded workflow decision under the tested conditions.

Sources checked

Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.