The numbers come from different tests
| Published figure | Named evaluation | Source named by Google | Valid use |
|---|---|---|---|
| 90.7% | Harvey LAB-AA | Artificial Analysis public leaderboard | Describe the 3.7 chart with its date and method. |
| 10.0% | Harvey’s Legal Agent Benchmark | Vals.AI | Compare only with values in the same 3.8 chart and setup. |
| 8.8% | The same 3.8-chart benchmark | The same chart provenance | A 10.0 versus 8.8 comparison is arithmetically compatible, subject to the disclosed setup. |
Neither figure is a universal legal-accuracy rate
A benchmark score does not show the probability that a model will answer a client’s question correctly. The result belongs to a dataset, harness, scoring method, model ID, sampling configuration and date. It does not automatically transfer to a firm’s authorities, documents, tools, jurisdiction or review process.
Read the methodology before the headline
Google states that Gemini scores are generally pass at 1 unless otherwise noted and that runs use the named API model with default sampling unless specified. It also says many non-Gemini results are provider-reported and several Gemini results are self-computed. Those facts belong beside any comparative claim.
What the benchmark can change
The 3.8 result can justify adding Gemini 3.8 Flash to a bounded regression test. It cannot justify a firm-wide switch. Use an internal set with controlling sources, expected issues and documented failure classes. Count unsupported authority, missed material issues, reviewer corrections, completion time and total cost.
Minimum benchmark record
- Exact benchmark name and version.
- Publisher and underlying result source.
- Dataset and scoring unit.
- Model ID, tools, sampling and attempt policy.
- Comparator configuration.
- Publication date and access date.
- Known transfer limits.
What remains Unknown
The public record does not establish Gemini 3.8 performance on your matters, the cause of a score difference, professional compliance, review burden or completed-task economics. Those require a disclosed local test.
FAQ
Did Gemini legal performance fall from 90.7% to 10.0%?
The public documents do not support that conclusion. The figures carry different benchmark names and underlying sources.
Can 10.0% be compared with 8.8%?
They can be described as a same-chart comparison when the model configurations and methodology shown with that chart are preserved.
Does a legal benchmark prove accuracy for a law firm?
No. It can inform a test shortlist, but the firm must validate its own documents, authorities, workflows and review burden.
Sources checked
- Gemini 3.8 Flash model card, checked 2026-09-03.
- Gemini 3.8 evaluation methodology, checked 2026-09-03.
- Gemini 3.7 evaluation methodology, checked 2026-09-03.
Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.