Download the scorecard
The template includes task identity, baseline and candidate versions, criteria counts, blocking failures, source accuracy, material omissions, reviewer corrections and minutes, run cost, latency, refusal and stop results, multilingual notes, rollback result, decision and reviewer. No sample row is labeled as evidence.
Six gates before a weighted score
| Gate | Minimum record | Blocking example |
|---|---|---|
| Frozen baseline | Product, account, model ID, prompt, sources, tools and date | The comparator changes mid-pilot. |
| Material quality | Expected issues, citations, omissions and prohibited assertions | Unsupported authority or missed material protection. |
| Reviewer effort | Qualified reviewer, corrections and minutes | No accountable final reviewer. |
| Completed-task cost | Model, tools, retries and reviewer time | Generated draft counted as completed work. |
| Permission and stop test | Allowed tools, destinations, denials and abstentions | Unapproved read, write or send succeeds. |
| Rollback evidence | Prior route, trigger, owner and tested recovery | The workflow cannot return to the approved baseline. |
Use explicit formulas
- Criterion pass rate = passed criteria ÷ evaluated criteria × 100.
- Strict task pass = 1 only when all blocking criteria and the defined minimum pass.
- Source accuracy = verified cited propositions ÷ checked cited propositions × 100.
- Completed-task cost = model + tools + retries + reviewer minutes ÷ 60 × loaded reviewer rate.
- Review delta = candidate reviewer minutes − baseline reviewer minutes.
Do not average away a blocking failure. A high criterion rate with an invented authority still fails the task.
Choose a small representative set
Include ordinary work, long material, conflicting sources, missing information, adversarial instructions, tool failure, required abstention and one edge case from the actual practice. For multilingual work, include the languages and jurisdictions the firm will use, not a generic translation prompt.
Keep the same source packet, instructions, output format, tools and reviewer standard across the baseline and candidate. If a condition changes, start a new comparison row.
Decision states
| State | Meaning |
|---|---|
| Approve bounded pilot | All blocking gates pass for the named workflow, data class, account and period. |
| Restrict | The route is useful only with narrower data, tools, users or review. |
| Revise and retest | A remediable control or configuration failed. |
| Reject | A material quality, permission, economics or rollback boundary fails. |
| Unknown | Required evidence was unavailable; absence of evidence is not a pass. |
What the score cannot establish
The scorecard supports a workflow decision. It does not certify the model, preserve privilege, prove compliance, predict all matters or replace legal judgment. ABA guidance and NIST can inform governance, but the controlling professional obligations and risk assessment depend on the facts and jurisdiction.
FAQ
What should a legal AI pilot measure?
Measure material quality, source accuracy, reviewer corrections and time, completed-task cost, latency, permissions, stops, multilingual fit and rollback against a frozen baseline.
Can a weighted score offset a serious legal failure?
No. Define blocking criteria first. Unsupported authority, prohibited disclosure or unauthorized action should fail the task regardless of the average.
Does a passing score certify the AI model?
No. It supports a dated, bounded workflow decision under the tested conditions.
Sources checked
- AI Risk Management Framework, checked 2026-09-03.
- Formal Opinion 512, checked 2026-09-03.
- Fable 5.1 and Mythos 5.1 system card, checked 2026-09-03.
- Gemini 3.8 evaluation methodology, checked 2026-09-03.
Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.