What Anthropic tested
| Field | System-card disclosure |
|---|---|
| Benchmark | Harvey AI Legal Agent Benchmark, an open-source benchmark spanning more than 1,200 tasks and 24 practice areas. |
| Evaluated set | 1,235 of 1,251 problems; 16 data-defect exclusions were identified before testing. |
| Rubric | Expert-written criteria per task: minimum 23, median 56, maximum 194. |
| Configuration | Adaptive thinking, max effort, public Messages API and production safeguards active. |
| Runs | Five runs for the reported public-set all-pass result. |
| Harness | Anthropic internal reimplementation with a reduced toolset and the default Claude Sonnet 4.6 judge. |
The two percentages answer different questions
| Metric | Question | Reported result |
|---|---|---|
| Mean criterion-pass | Across completed runs, what fraction of individual rubric conditions passed? | 90.81% |
| All-pass | For what fraction of tasks did every rubric criterion pass? | 19.09% ± 0.92, n=5 |
A task with 55 of 56 criteria satisfied can contribute a very high criterion-pass result and still count as a failure in the strict all-pass metric. The all-pass score is therefore closer to a complete acceptance gate, but it still belongs to this benchmark and harness.
Why neither number is legal accuracy
The benchmark uses specified documents, tasks, tools and an LLM judge. Anthropic also says its harness differs from the public harness and reports safeguards and grader-pipeline changes. Those details matter. The result does not show accuracy on a firm’s authorities, language, jurisdiction, templates, permissions or review process.
Label the result vendor-reported. Preserve the exact system-card date and configuration. Do not combine it with a score from a different held-out set, harness or effort level as if the denominators matched.
Build a task-level acceptance test
- Choose one repeatable task with an approved source packet and expected output.
- Write the material criteria before running the model.
- Mark a few criteria as blocking: unsupported authority, missed deadline, prohibited disclosure or unapproved external action.
- Run the current baseline and candidate with the same account, tools and constraints.
- Record every criterion, the strict task result, reviewer corrections, minutes and completed-task cost.
- Reject the route when any blocking criterion fails, regardless of the average score.
What the benchmark can support
The result can justify putting Fable 5.1 on a shortlist and asking better diligence questions. It cannot approve confidential data, establish professional compliance or predict savings. Procurement should connect the score to model availability, retention, contract, cloud route and a dated internal test.
FAQ
How can 90.81% and 19.09% both be correct?
The first averages individual rubric criteria. The second counts a task only when every criterion passes. One missed requirement can fail the whole task.
Did Anthropic use the public Harvey benchmark harness?
Anthropic says it used an internal reimplementation that preserves the tasks, rubric, all-pass scoring and default judge but exposes a reduced toolset.
Does 19.09% mean Fable is correct on one in five legal questions?
No. It is a strict task-level result for this benchmark, harness and configuration, not a universal legal-question accuracy rate.
Sources checked
- Fable 5.1 and Mythos 5.1 system card, checked 2026-09-03.
- Claude Fable 5.1 overview, checked 2026-09-03.
- AI benchmark procurement checklist, checked 2026-09-03.
Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.