Sonnet 5.5 API pricing: calculate cache and task costs
Calculate Claude Sonnet 5.5 API costs for input, output, caching and Batch jobs, with worked examples that separate token prices from completed-task costs.
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens on the standard Claude API. A cache read costs $0.20 per million tokens. Those are vendor list prices checked on September 29, 2026, not an Ofox quote or the price of a Claude subscription. The important budgeting question is how many billable tokens your application uses to finish an accepted task.
For developers moving from Sonnet 5, the rate card itself is unchanged. Anthropic’s claim of lower costs on much of its work concerns model behavior and token consumption; it is not a blanket discount on every request. Start with the official model specification, then estimate the workload you actually intend to run.
The Sonnet 5.5 rate card
| Token category | USD per million tokens | What to count |
|---|---|---|
| Uncached input | $2.00 | Input not charged as cache creation or reading |
| Output | $10.00 | Billable output, including billed thinking usage |
| Five-minute cache creation | $2.50 | Tokens written to that cache tier |
| One-hour cache creation | $4.00 | Tokens written to that cache tier |
| Cache reading | $0.20 | Tokens actually served from the prompt cache |
The model page also lists a 50% Batch API discount on input and output and a 512-token minimum cacheable prompt. Reaching that length does not automatically create a cache: configure caching and check the usage fields. Do not count cached tokens twice by adding them to ordinary input again.
Tool charges, data-residency options and other service-specific charges need separate checking against the full pricing documentation. A provider’s supported features, currency and billing rules can differ. Check its current terms before applying this vendor calculation to another endpoint.
A small request: $0.04, not a monthly budget
For a hypothetical request using 10,000 uncached input tokens and 2,000 billable output tokens:
Input: 10,000 / 1,000,000 × $2 = $0.02
Output: 2,000 / 1,000,000 × $10 = $0.02
Total: $0.04
At exactly that usage, 1,000 requests would cost $40 before other charges. This is arithmetic, not a measurement of Sonnet’s typical response length. A longer reasoning trace, another tool turn or a retry changes the total. Multiplying the visible answer’s word count by a guessed conversion factor is not a reliable billing method.
When caching changes the calculation
Suppose ten requests share a 100,000-token prefix. Assume one five-minute cache write, nine actual cache hits within the relevant lifetime, and no other input or output. The prefix costs $0.25 to write plus 9 × $0.02 to read: $0.43, versus $2.00 for ten fully uncached copies. That is a saving of $1.57, or 78.5%, for this prefix alone.
It is not a 78.5% discount on the whole application. Changing the prefix, missing the cache or generating substantial output can reduce the overall saving. Record cache creation and read usage for each request. A cache-shaped prompt is not evidence that the provider actually billed a cache hit.
For eligible non-interactive work, evaluate Batch separately. Applying its advertised 50% input/output discount to the uncached $0.04 example gives $0.02 under the stated assumptions. This example does not combine Batch with caching or assert how every combination is billed.
Measure accepted-task cost
Use this worksheet for a bug fix, document or extraction job:
Task ID | Model | Effort | Input | Cache write | Cache read
Output | Tool charges | Attempts | Accepted? | Total USD
Include failed attempts and retries in the cost of the task. Divide the sum by accepted tasks only after defining acceptance: tests pass, required fields are correct, or the deliverable meets a review checklist. Keep subscription usage in a separate column; a CLI’s estimated API-equivalent cost is not automatically a charge on your subscription invoice.
There is a reason to record effort. Artificial Analysis’s launch evaluation found very high output consumption at max effort. Its benchmark workload does not predict your bill. It also used a pre-release deployment affected by a structured-output issue, so its own rerun caveat matters. This is a warning to measure task usage, not a reason to call Sonnet universally expensive.
Choose the next check
If the rate card is clear but the bill is not, compare actual output tokens, cache hits and retry counts before switching providers. For tuning, use the Sonnet 5.5 effort guide. For model selection, see Sonnet versus Opus for coding tasks.
Ofox’s model catalog is a place to check currently offered models and provider details. This article does not establish that a particular Sonnet endpoint, price or feature is available there. Verify the exact model and protocol before sending a request.
Frequently Asked Questions
- Did Sonnet 5.5 lower the token price from Sonnet 5?
- No. The current official documentation says the prices are unchanged. A task can still use fewer or more tokens, which changes its total cost without changing the per-token rate.
- Does a Claude Pro or Team subscription include API credits?
- Do not treat a subscription allowance as API credit. Confirm the billing route used by your client and review the relevant plan separately from the Claude API rate card.
- Is max effort the cheapest way to finish a task?
- There is no general guarantee. Compare accepted results at a few effort settings, including retries and elapsed time. A higher setting can consume much more output without improving your acceptance rate.


