Benchmark provenance · source checked

Gemini 3.8 Legal Benchmark: Why 90.7% and 10.0% Do Not Trend

Google’s 90.7% and 10.0% legal figures look like a collapse only when their benchmark names, sources and scales are removed. The documents identify different evaluations.

Direct answer

Do not subtract Gemini 3.8’s 10.0% from Gemini 3.7’s 90.7%. Google’s methodology identifies the earlier figure as Harvey LAB-AA sourced from Artificial Analysis and the later figure as Harvey’s Legal Agent Benchmark sourced from Vals.AI. Only compare values shown for the same benchmark and chart, such as the reported 10.0% versus 8.8% within the 3.8 chart.

Two separate evidence dossiers showing why the Gemini 3.7 and 3.8 legal benchmark percentages are not one trend
Review frame. Only values inside the same disclosed chart can support a direct comparison.

The numbers come from different tests

Published figureNamed evaluationSource named by GoogleValid use
90.7%Harvey LAB-AAArtificial Analysis public leaderboardDescribe the 3.7 chart with its date and method.
10.0%Harvey’s Legal Agent BenchmarkVals.AICompare only with values in the same 3.8 chart and setup.
8.8%The same 3.8-chart benchmarkThe same chart provenanceA 10.0 versus 8.8 comparison is arithmetically compatible, subject to the disclosed setup.

Neither figure is a universal legal-accuracy rate

A benchmark score does not show the probability that a model will answer a client’s question correctly. The result belongs to a dataset, harness, scoring method, model ID, sampling configuration and date. It does not automatically transfer to a firm’s authorities, documents, tools, jurisdiction or review process.

Read the methodology before the headline

Google states that Gemini scores are generally pass at 1 unless otherwise noted and that runs use the named API model with default sampling unless specified. It also says many non-Gemini results are provider-reported and several Gemini results are self-computed. Those facts belong beside any comparative claim.

What the benchmark can change

The 3.8 result can justify adding Gemini 3.8 Flash to a bounded regression test. It cannot justify a firm-wide switch. Use an internal set with controlling sources, expected issues and documented failure classes. Count unsupported authority, missed material issues, reviewer corrections, completion time and total cost.

Minimum benchmark record

  1. Exact benchmark name and version.
  2. Publisher and underlying result source.
  3. Dataset and scoring unit.
  4. Model ID, tools, sampling and attempt policy.
  5. Comparator configuration.
  6. Publication date and access date.
  7. Known transfer limits.

What remains Unknown

The public record does not establish Gemini 3.8 performance on your matters, the cause of a score difference, professional compliance, review burden or completed-task economics. Those require a disclosed local test.

FAQ

Did Gemini legal performance fall from 90.7% to 10.0%?

The public documents do not support that conclusion. The figures carry different benchmark names and underlying sources.

Can 10.0% be compared with 8.8%?

They can be described as a same-chart comparison when the model configurations and methodology shown with that chart are preserved.

Does a legal benchmark prove accuracy for a law firm?

No. It can inform a test shortlist, but the firm must validate its own documents, authorities, workflows and review burden.

Sources checked

Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.