Anthropic launched Sonnet 5.5 on September 28. Its current documentation establishes the product and technical specifications. Whether it improves your legal work remains a question for a controlled test. Anthropic release notes
For a managing partner, the immediate decision is where to trial it. For legal operations, the harder decision is which account, data path and tools the trial should use. A model choice alone leaves most of that work unfinished.
What changed with Sonnet 5.5
Anthropic presents Sonnet 5.5 as a faster, lower-cost option for well-scoped everyday tasks, including document creation. It reports efficiency gains over Sonnet 5 and reserves its stronger recommendation for Opus 5.5 on complex, open-ended work. Treat that positioning as a reason to investigate specific tasks. It is not a legal-workload allocation. Launch announcement
The Claude Platform lists a 1M-token context window, a 128K-token maximum output and the API model ID claude-sonnet-5-5. Standard API input and output rates are $2 and $10 per million tokens. Those specifications describe capacity and billing; they do not show that a long agreement will be read completely or that its difficult provisions will be identified. Model specifications
The useful question is narrower: can the model produce an output that your reviewer can verify with fewer corrections or less effort than the current method?
Three legal tasks worth testing first
These are proposed evaluations, not reported Sonnet 5.5 results.
Compare a clause with an approved playbook. Supply the clause, its definitions and the relevant playbook rules. Ask for each deviation, the supporting text and a proposed edit. Keep the reviewer responsible for whether the edit fits the transaction. A fluent replacement clause earns little if it quietly changes an agreed exception.
Extract obligations from a supplied agreement. Request the responsible party, action, trigger, deadline and source location. Require the output to identify missing or ambiguous information rather than complete it from habit. The reviewer should be able to move from every row back to the agreement. An attractive obligations table with a wrong notice trigger is still wrong.
Draft from verified notes. Give the model the facts and audience for a status update or internal summary. Ask it to preserve uncertainty and avoid adding commitments. Review names, amounts, dates and anything that sounds like a promise before the draft leaves the team.
These tasks have an advantage for evaluation: you can define a useful output and inspect the source material. Open-ended legal research introduces a different job. The team must establish whether each authority exists, supports the proposition, applies in the relevant jurisdiction and remains current. Keep that research evaluation separate from a test of drafting style.
Choose the deployment before the documents
“We use Claude” is too broad for a matter record. Write down whether the work uses a Claude application, the direct API, a cloud provider or a legal product that incorporates Claude. Then identify the account, agreement, enabled features and person responsible for the deployment. The Claude product and data guide explains those layers.
Use three questions to make the choice concrete:
Where will the input, output and working files be stored?
Which tools or external services can receive information during the task?
Who can review, export, delete or stop the work?

Deployment decision map. These are checks for the selected route, not a certification or an exhaustive list of products. API ZDR eligibility and third-party data paths require separate review. Retention scope and feature eligibility, checked October 2, 2026.
For Sonnet 5.5, a particularly important distinction is zero data retention. The model can be used with ZDR, but the applicable arrangement depends on the surface and features. Anthropic’s API documentation excludes ordinary Team and Enterprise interfaces from that API arrangement; qualified Enterprise Claude Code access has a separate offering. External integrations are also outside its scope. Check feature eligibility and the operative agreement rather than copying a model-level label into a policy. API retention documentation
Make the tool decision separately. An approved document-only comparison does not require permission to search the entire document-management system or send the finished draft. Add a connector when the workflow needs it and the recipient, access and retention questions have answers.
Professional obligations need their own review. Anthropic has published a legal-use configuration discussion, but it is the vendor’s perspective, not a ruling on your matter. Have the appropriate firm owner assess client requirements, confidentiality, privilege and jurisdiction before using matter material. A retention setting cannot make that decision for them.
A contract example with a reviewable result
Consider a synthetic evaluation involving a sample services agreement. The playbook says renewal requires written agreement. The sample instead provides automatic renewal unless notice arrives before a specified deadline. It also contains an exception in a schedule that affects one service.
Ask Sonnet to produce a short deviation record with six fields: clause location, relevant wording, playbook rule, deviation, proposed edit and unresolved question. Require it to review the schedule before concluding that the renewal rule applies throughout the agreement.
The expected answer key belongs to the reviewer. It should identify the renewal conflict, the notice mechanism and the schedule exception. The model must not invent a deadline, omit the exception or label the proposed edit as already agreed.
Now run the same packet through the approved baseline. Record missed issues and material corrections before comparing readability. Time the review, including the work needed to open the cited clauses. If Sonnet produces a smoother summary but hides the exception, the trial has exposed a failure. If it produces a useful deviation record, that supports continued testing of this task under these conditions.
This example is a proposed protocol. AI Vortex has not run it on Sonnet 5.5 or measured its results.
Measure completed work rather than cheap tokens
API prices and subscription charges answer different purchasing questions. For an API workflow, capture the actual model and tool bill. For a seat-based product, understand its allowance and usage terms. In both cases, include retries and the review needed to make the output usable.
A simple internal measure is:
Completed-task cost = model and tool charges + retry charges + reviewer time at the firm’s chosen internal rate.
Keep the rate and assumptions visible. Do not count the same retry or review time twice. This is an evaluation measure, not a client-billing recommendation.
The legal AI pilot scorecard already provides a place to record quality, reviewer effort, costs, permissions and rollback. Use it for a representative sample, including difficult documents and known exceptions. A single successful demonstration cannot establish the share of a practice group’s work that should move to Sonnet.
Sonnet or Opus needs a task-level answer
Start with the same source packet and scoring rules. Test a routine task and a genuinely difficult version of it. Keep the account, tools and effort settings in the record so another reviewer can understand the comparison.
Choose Sonnet for a workflow when its observed results meet the quality threshold and its review burden and cost make sense. Choose another route when it fails that threshold. Raising the model tier should not remove the reviewer or weaken the acceptance criteria.
General benchmarks can help decide what deserves a test. They cannot tell a firm how often a model will miss a negotiated exception or require substantive correction on its matters. The benchmark procurement checklist helps separate a reported score from the evidence needed for adoption.
Record the model that actually answered
Some flagged Sonnet 5.5 requests can fall back to Sonnet 5. In the API, automatic switching is not enabled by default; customers must opt in and configure fallbacks. Anthropic says the app displays the switch, and material read from files or connectors can trigger checks. A trial record should therefore capture the responding model, any refusal and any fallback, rather than just the model selected at the start. This does not establish that routine legal tasks commonly trigger switching. Model-switching guidance
If the team operates a custom API application, give its developer the current migration guide. Existing thinking settings and response parsing can require changes; a model-name substitution is not a complete regression test. Check extraction, export and failure handling before changing the production default. API migration guide
Make the next decision small enough to inspect
Choose one repeatable task, one approved data set and one accountable reviewer. Write down the material failures that would stop adoption. Preserve the current route and test that the team can return to it.
Then evaluate Sonnet 5.5 on the work itself. The result should be a decision you can explain: this task is approved under these conditions, this one needs more testing, and this one stays with the existing process.
Operational information, not legal advice. Product facts were checked October 2, 2026; verify access, pricing and terms again before procurement or deployment.
The contract playbook ownership guide explains who approves the standard used in the sample evaluation. The hybrid-firm essay puts completed-task economics in the wider delivery context.
Questions and answers
When did Anthropic release Claude Sonnet 5.5?
Anthropic launched Claude Sonnet 5.5 on September 28, 2026. Vendor positioning and benchmarks can inform which tasks to test; they do not establish legal-workload suitability.
What are the standard Sonnet 5.5 API rates?
As checked October 3, 2026, standard API rates are $2 per million input tokens and $10 per million output tokens. Completed-task cost also includes tools, retries and reviewer time.
Is Sonnet 5.5 compatible with zero data retention?
The model is compatible with ZDR, but coverage depends on the deployment, features and operative arrangement. Do not extend API terms to ordinary Claude applications or external integrations.
Which legal tasks should a team test first?
Candidate tests include a clause comparison against an approved playbook, obligations extraction from a supplied agreement and drafting from verified notes. These are proposed evaluations, not reported Sonnet results.
How much legal work should move to Sonnet?
That share is unknown without a controlled evaluation of the firm’s actual tasks. Compare material corrections, source fidelity, reviewer time and completed-task cost before assigning a workflow.
Operational information, not legal advice. Product details and primary sources rechecked October 3, 2026. Proposed evaluations are identified in the text.