Benchmark interpretation · source card checked

Claude Fable 5.1 Legal Benchmark: 90.81% vs 19.09%

Anthropic reports two correct but very different views of the same test. One averages rubric criteria. The other asks whether every criterion in a task passed.

Direct answer

Anthropic reports that Fable 5.1 met 90.81% of individual rubric criteria on average but passed every criterion on 19.09% of evaluated tasks. The gap is expected when each task contains 23 to 194 criteria and one missed requirement makes the task fail under the strict score. Both figures are vendor-reported results from Anthropic’s internal reimplementation, not a law-firm accuracy rate.

Evidence fields separating mean criterion-pass rate from strict all-pass task rate in the Fable 5.1 legal benchmark
Review frame. A high average across criteria can coexist with a much lower rate of tasks that pass every criterion.

What Anthropic tested

FieldSystem-card disclosure
BenchmarkHarvey AI Legal Agent Benchmark, an open-source benchmark spanning more than 1,200 tasks and 24 practice areas.
Evaluated set1,235 of 1,251 problems; 16 data-defect exclusions were identified before testing.
RubricExpert-written criteria per task: minimum 23, median 56, maximum 194.
ConfigurationAdaptive thinking, max effort, public Messages API and production safeguards active.
RunsFive runs for the reported public-set all-pass result.
HarnessAnthropic internal reimplementation with a reduced toolset and the default Claude Sonnet 4.6 judge.

The two percentages answer different questions

MetricQuestionReported result
Mean criterion-passAcross completed runs, what fraction of individual rubric conditions passed?90.81%
All-passFor what fraction of tasks did every rubric criterion pass?19.09% ± 0.92, n=5

A task with 55 of 56 criteria satisfied can contribute a very high criterion-pass result and still count as a failure in the strict all-pass metric. The all-pass score is therefore closer to a complete acceptance gate, but it still belongs to this benchmark and harness.

Why neither number is legal accuracy

The benchmark uses specified documents, tasks, tools and an LLM judge. Anthropic also says its harness differs from the public harness and reports safeguards and grader-pipeline changes. Those details matter. The result does not show accuracy on a firm’s authorities, language, jurisdiction, templates, permissions or review process.

Label the result vendor-reported. Preserve the exact system-card date and configuration. Do not combine it with a score from a different held-out set, harness or effort level as if the denominators matched.

Build a task-level acceptance test

  1. Choose one repeatable task with an approved source packet and expected output.
  2. Write the material criteria before running the model.
  3. Mark a few criteria as blocking: unsupported authority, missed deadline, prohibited disclosure or unapproved external action.
  4. Run the current baseline and candidate with the same account, tools and constraints.
  5. Record every criterion, the strict task result, reviewer corrections, minutes and completed-task cost.
  6. Reject the route when any blocking criterion fails, regardless of the average score.

What the benchmark can support

The result can justify putting Fable 5.1 on a shortlist and asking better diligence questions. It cannot approve confidential data, establish professional compliance or predict savings. Procurement should connect the score to model availability, retention, contract, cloud route and a dated internal test.

FAQ

How can 90.81% and 19.09% both be correct?

The first averages individual rubric criteria. The second counts a task only when every criterion passes. One missed requirement can fail the whole task.

Did Anthropic use the public Harvey benchmark harness?

Anthropic says it used an internal reimplementation that preserves the tasks, rubric, all-pass scoring and default judge but exposes a reduced toolset.

Does 19.09% mean Fable is correct on one in five legal questions?

No. It is a strict task-level result for this benchmark, harness and configuration, not a universal legal-question accuracy rate.

Sources checked

Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.