Sonnet 5.5 effort: when to use medium, high or max
Choose Sonnet 5.5 effort using acceptance tests, latency and task cost. Learn why API and Claude Code defaults differ and when higher effort is worth testing.
Use Sonnet 5.5 effort as a parameter to evaluate, not a quality guarantee. Anthropic recommends starting at medium for well-specified agentic coding and multistep tool tasks, moving to high for harder or longer work. For latency-sensitive chat it suggests medium or low. The native API’s default remains high; Claude Code has its own default.
Those recommendations come from the Sonnet 5.5 behavior documentation, checked September 29, 2026. This article explains how to test the trade-off. It does not claim an Ofox benchmark across all effort levels.
Start with the work, then choose a setting
| Workload | Documented starting direction | What to inspect |
|---|---|---|
| Short, latency-sensitive chat | low or medium | Response time and missing required details |
| Well-specified agentic coding | medium | Tests, scope and tool turns |
| Harder or longer tool tasks | high | Accepted result and repeated failure modes |
| A difficult task still failing | Test higher effort against a baseline | Whether extra cost changes the outcome |
The last row is an evaluation suggestion, not an official promise that xhigh or max will fix the failure. A task with missing requirements, unavailable tools or contradictory instructions may remain unsolved at any effort.
The model page specifies high as the API default. Claude Code’s configuration documentation describes medium for Sonnet 5.5 in that client. Record the entry point before comparing results. Otherwise two runs described as “default Sonnet” may use different settings.
Why max is not a universal recommendation
Artificial Analysis’s launch evaluation found high output-token consumption at max effort and an unfavorable cost trade-off against some alternatives. The same report shows substantial benchmark capability. Both can be true: a model can reach a strong result by spending more tokens.
That report tested a pre-release deployment with a structured-output bug and states that relevant evaluations will be rerun. Its benchmark cost is not a price quote for your bug fix, nor does it establish that every max run is wasteful. Use it as a reason to collect cost and quality together.
Higher effort can also change the kind of behavior you see. More exploration is useful when it tests relevant explanations; it is harmful when the model expands beyond the requested change or spends time on irrelevant work. Your acceptance criteria should include scope, not just whether one test passes.
Design a small effort sweep
Prepare a set of representative tasks with expected outputs or acceptance checks. Keep the model version, tools, inputs and starting repository state fixed. Run the baseline setting, then test a higher or lower level on the same tasks. Use independent sessions when you want to avoid the first run teaching the next run the answer.
Record at least:
Task | Effort | Applied setting | Accepted | Attempts
Elapsed seconds | Input tokens | Cache categories | Output tokens
Tool charges | Total cost | Out-of-scope edits | Review notes
Separate first-attempt results from results after retries. If you cap time or tokens, mark a capped run explicitly rather than interpreting it as an ordinary completed answer. Preserve failures in the dataset. A table containing only successful examples cannot establish the most reliable configuration.
One useful decision rule is to retain a higher setting only when it improves an outcome that matters enough to justify its added cost or delay. Define that threshold before examining the results. For example, a team might value fewer incorrect patches more than slightly lower latency; a chat product might have the opposite constraint. Do not borrow a universal threshold from someone else’s benchmark.
Set effort without creating an invalid request
For the native API, an adaptive-thinking request can include:
{
"model": "claude-sonnet-5-5",
"max_tokens": 2048,
"thinking": {"type": "adaptive"},
"output_config": {"effort": "high"},
"messages": [{"role": "user", "content": "List the acceptance checks for a CSV parser fix."}]
}
This is a documentation-based request body, not a live API test. max_tokens caps thinking plus response text; thinking tokens are billed as output even when their text is omitted. It is not a requested amount of thinking or an all-inclusive dollar budget. Authentication, version headers and your response handling are separate requirements.
If you use between_tools to disable up-front thinking, keep effort at high or below. That mode does not support xhigh or max. Changing its effort mid-conversation also has constraints. Use the migration checklist before copying older disabled or manual-budget configurations.
In Claude Code, launch with --effort medium or select a supported level with /effort. Managed settings can cap what actually runs. Requested and applied effort must not be treated as equivalent without checking the client and account behavior.
Decide whether to change model instead
When a task remains difficult, compare increasing Sonnet effort with trying another model. Sonnet versus Opus addresses the within-Claude choice; Sonnet versus Sol covers a cross-provider trial. Keep the task and acceptance test fixed when changing the configuration.
Once you choose a baseline, record why you chose it and which failures justify escalation. This makes future model updates easier to evaluate. Sonnet 5.5’s effort levels were recalibrated relative to Sonnet 5, so reusing the old label without testing is not evidence of equivalent behavior.
Frequently Asked Questions
- Is high the default everywhere?
- No. The native API and Claude Code have different documented defaults. Account controls and explicit settings can change what is applied.
- Can I use max with between_tools?
- No. The documented mode supports low, medium and high. Use adaptive thinking for higher effort levels.
- Will lower effort always reduce the cost of a completed task?
- Not necessarily. It may reduce tokens per attempt but require more attempts or fail more often. Measure accepted-task cost rather than only one response.


