Sonnet 5.5 vs Opus 5.5: which coding tasks need Opus?

Choose Sonnet 5.5 or Opus 5.5 for bug fixes, code review and complex changes. Compare task scope, effort, cache costs and the evidence behind the decision.

Art-line illustration of hands with the title Sonnet 5.5 vs Opus 5.5.

Start by evaluating Sonnet 5.5 on well-defined coding tasks and Opus 5.5 on work where judgment, ambiguity or repeated failure justifies the extra trial. That is a workload-selection approach, not proof that Sonnet always handles small tasks or that Opus always wins on large ones. Your repository tests and review standards remain the acceptance criteria.

Anthropic positions Sonnet as a faster, lower-cost complement to Opus, while describing Opus as suited to complex work requiring careful judgment. The two models’ billing and effort settings make the choice less simple than “Sonnet is half the price.” This guide uses documentation checked September 29, 2026; no original head-to-head model trial is claimed.

Rates do not settle the choice

Official Claude API categorySonnet 5.5Opus 5.5
Input, per million tokens$2$4
Output, per million tokens$10$20
Cache reads, per million tokens$0.20$0.20
API default efforthighmedium
Context window1M tokens1M tokens

Sources: Sonnet specification and Opus specification. These are vendor rates, not subscription quotas or an Ofox quote. The cache-read row is especially relevant to repeated repository context: a large cached prefix does not have the same two-to-one difference as uncached input and output.

Imagine two synthetic requests with the same 100,000 cache-read tokens and 2,000 billable output tokens, with all other billable categories excluded. Sonnet would cost $0.04 and Opus $0.06. The difference is not twofold because cache reading costs the same in this example. If one model takes additional turns, the comparison changes again.

Match the task to a testable outcome

TaskUseful starting comparisonEvidence to retain
Bug with a deterministic reproductionSonnet at a modest effort setting, then Opus if neededFailing test, patch, full regression run
Small feature with explicit requirementsSonnet versus your current working baselineAcceptance checklist, scope of edits
Ambiguous cross-module failureInclude Opus from the outsetCompeting explanations, inspected files, verified cause
Repository reviewGive both the same scopeConfirmed findings versus false positives
Risky migrationCompare planning and validation separatelyMigration checklist, rollback path, integration tests

These are trial designs. They are not measured pass-rate claims. A short patch can require difficult reasoning, and a long but mechanical change can be easy. File count or lines changed alone is a poor measure of task difficulty.

Read benchmark claims with their conditions

Artificial Analysis’s Sonnet launch report shows strong results on several tasks but much higher output consumption at max effort. It also notes a pre-release structured-output issue affecting the tested deployment and planned reruns. Those results can justify testing both models; they cannot establish your cheapest configuration.

Do not compare an Opus medium run to a Sonnet max run and call the result a pure model comparison. It is a configuration comparison. That can still be valuable, but publish both settings and the actual budgets. Equal effort names also do not guarantee equal computation.

Anthropic’s own release announcement describes different model strengths and evaluation conditions. Keep vendor claims separate from independent measurements and from your own observations. A disagreement between charts may reflect tasks or settings rather than a mysterious contradiction.

A simple escalation policy

Before a run, write a stopping condition: for example, one proposed fix plus a regression check. If the check fails, inspect the failure before rerunning. If the model misunderstood the task, clarify the input rather than blindly increasing effort. If it found the right area but cannot produce a valid fix, trying Opus becomes a useful controlled escalation.

Carry the issue description, relevant files and test results into the new run. Do not assume encrypted thinking blocks are portable across models. Sonnet 5.5 has model- and conversation-specific rules described in the migration guide. Retain visible evidence rather than relying on hidden reasoning continuity.

Record the cost of the initial attempt and the escalation together. Otherwise a routing system can appear cheaper by attributing the first failure to one model and counting only the final successful run for another. Also record review time: a patch that compiles but needs substantial cleanup has not finished the job.

What to do on a subscription

Do not translate the API table into an exact number of Claude Code prompts. Subscription allowances and usage limits depend on the account and service conditions. Check the model actually selected, especially after a client update or provider change, and keep billing mode separate from the model name.

The Claude Code setup guide explains version checks and explicit selection. The effort guide separates API defaults from Claude Code defaults. If you are considering another provider, the Sonnet versus Sol guide gives a controlled comparison method.

Frequently Asked Questions

Is Sonnet always half the cost of Opus?
No. Its listed uncached input/output rates are half, but cache rates, token consumption, retries and tools affect completed-task cost. The worked example above deliberately isolates those differences.
Should code review always use Opus?
No universal rule follows from the documentation. Compare confirmed findings and false positives on representative reviews, including the time a human spends checking each finding.
Can I move the same conversation between models?
Visible messages can be part of a supported workflow, but thinking blocks have compatibility rules. Check the migration documentation rather than assuming the entire hidden state transfers.