Claude Sonnet API pricing: Anthropic, September 2026
Sonnet 5 is the cheapest Sonnet on Anthropic's API at $2 input and $10 output per million tokens, two-thirds of Sonnet 4.6 ($3 and $15) on every rate. Its newer tokenizer counts about 30% more tokens, so it saves about 13% on the same text.

Claude Sonnet 5 is the cheapest Sonnet on Anthropic’s API, at $2 per million input tokens and $10 per million output tokens.1
Claude Sonnet 4.6 costs $3 input and $15 output, the same as Sonnet 4.5 and Sonnet 4.1
Those are Anthropic’s list prices in its pricing documentation, checked 28 September 2026. The headline gap is a third, but Sonnet 5 counts more tokens for the same text, so the real saving is smaller.
TL;DR
- Sonnet 5 undercuts every older Sonnet by exactly a third on every line: $2 input, $10 output, $2.50 for a 5-minute cache write, $4 for a 1-hour write and $0.20 for a cache hit, per million tokens.1
- Sonnet 4.6, 4.5 and 4 share one rate card: $3, $15, $3.75, $6 and $0.30.1
- Batch halves input and output: $1 and $5 on Sonnet 5, $1.50 and $7.50 on the older three.1
- Long context costs nothing extra on Sonnet 4.6 and Sonnet 5. A 900k-token request pays the same per-token rate as a 9k-token one. Anthropic publishes no such statement for Sonnet 4.5 or Sonnet 4.1
- For the same text, Sonnet 5 saves about 13%, not 33%. Its newer tokenizer produces roughly 30% more tokens.1
- Caching moves the bill more than the model choice. In our worked support-assistant example, caching cuts the daily cost by 77% on either model.
How this ranking was built
This ranking sits in our model infrastructure coverage and follows the publication’s research standard for captured and computed figures.
- Source: Anthropic’s pricing documentation at docs.claude.com, captured 28 September 2026 at 15:05 AEST. Anthropic’s summary pricing page, captured the same minute, is the cross-check. Neither page shows a release date, so the checked date stands in for one.
- Field: input price per million tokens, the documentation’s Base tokens: Input column. The page writes the unit as “MTok”.1
- Currency: the documentation states that all prices are in USD.1
- Tax: the documentation makes no tax statement for direct API billing.
- Inclusion: every model named Sonnet in the documentation’s model pricing table. The table holds 18 Claude models; 4 are Sonnet models, and all 4 are ranked. No merges.
- Cross-check: the summary page lists Sonnet 5 under “Latest models” and Sonnet 4.6 and Sonnet 4.5 under “Legacy models”. Every Sonnet figure it shows matches the documentation. It doesn’t list Sonnet 4.2
- Exclusions: subscription plans, per-search charges and the US-only multiplier. The body sections cover the multiplier.
- Ties: tied models share a rank and appear newest first.
Output prices give the same order, because every Sonnet’s output price is five times its input price.
Claude Sonnet models ranked by input price
| Rank | Model | Announced | Summary page status | Input price (US$ per million tokens) |
|---|---|---|---|---|
| 1 | Claude Sonnet 5 | 30 Jun 2026 | Latest | 2 |
| 2= | Claude Sonnet 4.6 | 17 Feb 2026 | Legacy | 3 |
| 2= | Claude Sonnet 4.5 | 29 Sep 2025 | Legacy | 3 |
| 2= | Claude Sonnet 4 | 22 May 2025 | Not listed | 3 |
Prices exactly as Anthropic’s model pricing table lists them; announcement dates from Anthropic’s Sonnet page.1,3
Three Sonnet releases, from May 2025 to February 2026, share one price on the current rate card. Sonnet 5 is the only one priced lower.
Anthropic’s summary page doesn’t list Sonnet 4, and neither page says whether it stays open to new API accounts. Confirm availability before planning on it.
What does each Sonnet model cost per million tokens?
Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. Sonnet 4.6, Sonnet 4.5 and Sonnet 4 each cost $3 and $15.1
| Model | Input | Output | 5-min cache write | 1-hour cache write | Cache hit | Batch input | Batch output |
|---|---|---|---|---|---|---|---|
| Sonnet 5 | $2 | $10 | $2.50 | $4 | $0.20 | $1 | $5 |
| Sonnet 4.6 | $3 | $15 | $3.75 | $6 | $0.30 | $1.50 | $7.50 |
| Sonnet 4.5 | $3 | $15 | $3.75 | $6 | $0.30 | $1.50 | $7.50 |
| Sonnet 4 | $3 | $15 | $3.75 | $6 | $0.30 | $1.50 | $7.50 |
US dollars per million tokens. The cache-hit column is the documentation’s Hits and refreshes.1
Every Sonnet 5 figure is two-thirds of the matching Sonnet 4.6 figure. No line gets a bigger or smaller cut.
Same rate doesn’t mean same bill. Anthropic says Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, and that Claude Sonnet 4.6 and earlier models use the previous one. The exact increase depends on the content and workload.1
Sonnet 5 comes after 4.7, so it counts on the newer tokenizer. Priced per unit of text, its input rate works out at about $2.60 against Sonnet 4.6’s $3, and its output at about $13 against $15. That’s roughly 13% cheaper, not 33%.
Tool definitions add fixed input too, and here Sonnet 5 is leaner. Any tool brings a tool-use system prompt of 354 tokens on Sonnet 5 with tool choice set to auto or none, against 497 tokens on Sonnet 4.6.1
Anthropic publishes no fast mode rate for Sonnet. Its fast mode pricing covers Opus 5.5, Opus 5 and Opus 4.8 only.1
What do prompt caching and batch processing cost on Sonnet?
On Sonnet 5, a 5-minute cache write costs $2.50 per million tokens, a 1-hour write $4 and a cache hit $0.20, against $2 for ordinary input. On the three older Sonnets the same lines cost $3.75, $6 and $0.30.1
Those prices follow fixed multipliers on each model’s input rate: 1.25 times for a 5-minute write, 2 times for a 1-hour write and 0.1 times for a hit.1
Anthropic’s documentation says a cache pays off after one read at the 5-minute duration, or after two reads at the 1-hour duration.1
On Sonnet 5, the 5-minute write costs $0.50 per million tokens more than plain input, and each hit saves $1.80. The 1-hour write costs $2 more, so it needs a second hit to come out ahead.
The constraint is the cache lifetime. A prefix that isn’t read again within five minutes has to be written again at the premium rate. That’s the case for paying for the 1-hour write.
The Batch API takes 50% off both input and output tokens.1
That puts Sonnet 5 at $1 input and $5 output, and Sonnet 4.6, Sonnet 4.5 and Sonnet 4 at $1.50 and $7.50.1
The discounts combine. Anthropic states that caching multipliers stack with other pricing modifiers, including the Batch API discount and data residency.1
Stacked, a cache hit inside a batch costs $0.10 per million tokens on Sonnet 5 and $0.15 on the older three. That’s our arithmetic from the published multipliers, not a figure Anthropic prints.
US-only inference goes the other way. For Claude 4.6 and later models, the inference_geo parameter set to US-only adds a 1.1x multiplier to input, output, cache writes and cache reads.1
On Sonnet 5 that means $2.20 input, $11 output and $0.22 per cache hit. On Sonnet 4.6 it means $3.30, $16.50 and $0.33.
Sonnet 4.5 and Sonnet 4 can’t be pinned to the US on Anthropic’s API. Earlier models don’t support the parameter, always use standard pricing, and return a 400 error if a request includes it.1
Does Sonnet charge more for long context?
No, not on Sonnet 4.6 or Sonnet 5. Anthropic says Claude 4.6 and later models include the full 1M-token context window at standard pricing.1
A 900k-token request is billed at the same per-token rate as a 9k-token request, and caching and batch discounts apply at standard rates across the full window.1
So a 900,000-token prompt costs $1.80 of input on Sonnet 5 and $2.70 on Sonnet 4.6. Served as a cache hit, the same prompt costs $0.18 and $0.27.
Anthropic’s Sonnet page describes Sonnet 4.6 as featuring a 1M context window.3
The captured documentation publishes no long-context rate for Sonnet 4.5 or Sonnet 4, and its flat-rate statement covers only Claude 4.6 and later. Confirm the context limit and any long-context charge in Anthropic’s current documentation before sending a large prompt to either.
What does a real workload cost on Sonnet 5 versus Sonnet 4.6?
In our two worked examples, Sonnet 5 costs $1,040 against $1,200 on Sonnet 4.6 for a batch-sized classification job, and about $143 against $165 for a day of a cached support assistant. Both results allow for the newer tokenizer. Both are illustrations built on Anthropic’s published rates, checked 28 September 2026, not measured bills.
Example 1: classifying 100,000 documents. Each document is 3,000 input tokens and each answer 200 output tokens, counted on Sonnet 4.6’s tokenizer. At Anthropic’s rough estimate of 0.75 words per token, that’s about 2,250 words a document.1
| Model | Standard price | Batch price |
|---|---|---|
| Sonnet 4.6, 4.5 or 4 | $1,200 | $600 |
| Sonnet 5, same token counts | $800 | $400 |
| Sonnet 5, 30% more tokens | $1,040 | $520 |
US dollars, computed from Anthropic’s per-million-token rates. 300 million input and 20 million output tokens on the older tokenizer.
A job like this doesn’t need an answer inside a second. Batch halves it on any Sonnet, which beats switching model.
Example 2: a support assistant for one day. It serves 10,000 requests. Each re-reads a cached 20,000-token system prompt and reference set, adds 1,000 fresh input tokens and returns 500 output tokens. The cache stays warm, so one 20,000-token write covers the day.
| Line | Sonnet 4.6 | Sonnet 5, same tokens |
|---|---|---|
| Cache hits (200M tokens) | $60.00 | $40.00 |
| Fresh input (10M tokens) | $30.00 | $20.00 |
| Output (5M tokens) | $75.00 | $50.00 |
| One 5-minute cache write (20,000 tokens) | $0.08 | $0.05 |
| Day total | $165.08 | $110.05 |
| Same day with no caching | $705.00 | $470.00 |
US dollars, computed from Anthropic’s per-million-token rates.
Caching cuts the day by 77% on either model. Moving from Sonnet 4.6 to Sonnet 5 cuts it by a third at identical token counts, or to about $143 once the 30% tokenizer uplift is applied.
The estimate holds token counts fixed, and that’s its limit. A model that needs fewer attempts or shorter answers changes the result more than the rate card does. For the other drivers of a bill, see our explainer on what controls AI inference cost.
Use one real workflow to test the claim. Log cache hits, cache writes, fresh input and output as separate lines for one day, then price each line at the rates above.
How does Sonnet compare with Opus and Haiku?
Sonnet 5 costs half of Opus 5.5 and twice Haiku 4.5 per input and output token. Opus 5.5 costs $4 and $20 per million tokens and Haiku 4.5 costs $1 and $5.1
| Model | Input | Output | 5-min cache write | Cache hit |
|---|---|---|---|---|
| Fable 5.1 | $10 | $50 | $12.50 | $0.25 |
| Opus 5.5 | $4 | $20 | $5 | $0.20 |
| Opus 4.6 | $5 | $25 | $6.25 | $0.50 |
| Sonnet 5 | $2 | $10 | $2.50 | $0.20 |
| Sonnet 4.6 | $3 | $15 | $3.75 | $0.30 |
| Haiku 4.5 | $1 | $5 | $1.25 | $0.10 |
US dollars per million tokens, from Anthropic’s model pricing table.1
Cache hits break the usual order. Sonnet 5 and Opus 5.5 both charge $0.20 per million cached tokens. Sonnet 4.6 charges $0.30, half as much again as Opus 5.5.
The decision changes with the share of a workload that re-reads cached context. For a cache-heavy agent, staying on Sonnet 4.6 pays more for its largest line than moving up to Opus 5.5 would. For fresh text and long answers, Sonnet 4.6 still costs three-quarters of Opus 5.5.
Anthropic’s guidance is to choose Haiku for simple tasks, Sonnet for most production workloads and Opus for the most complex reasoning.1
The full Claude rate card, including Opus, Fable, Mythos and Haiku, is in our Claude API pricing ranking.
Why do other published Sonnet prices disagree?
Most published Sonnet prices that differ from Anthropic’s come from four causes: a single cache-write figure, the US-only multiplier, partner-cloud pricing and pages written before Sonnet 5.
One cache-write figure. TypingMind’s calculator lists Sonnet 4.6 at $3.00 input, $15.00 output, $0.3000 cache read and $3.75 cache write per million tokens.4
That matches Anthropic’s 5-minute write. Anthropic’s own summary page shows one write figure too, and notes that its caching prices reflect the 5-minute duration.2
The 1-hour write, $6 on Sonnet 4.6, appears only in the documentation.1
The US-only multiplier. Some price trackers show Sonnet 4.6 at $3.00 to $3.30 per million input tokens. $3.30 is exactly 1.1 times $3, the US-only rate.
Partner clouds. Anthropic says Bedrock and Google Cloud have independent regional pricing.1
Its own marketplace channels don’t reprice. Claude Platform on AWS rates usage at standard per-model rates in US dollars, applies any negotiated discount, then converts to Claude Consumption Units at $0.01 each.1
Dates. Sonnet 5 was announced on 30 June 2026.3
A page written before then quotes Sonnet 4.6 at $3 and $15 as the current Sonnet. Those numbers are still right for Sonnet 4.6. They’re no longer the cheapest Sonnet.
What changed since Sonnet 4.6?
Sonnet 5 cut every Sonnet rate by a third, from $3 input and $15 output to $2 and $10 per million tokens. Cache hits fell from $0.30 to $0.20 and batch input from $1.50 to $1.1
Anthropic’s Sonnet page states the same $2 and $10, with up to 90% savings from prompt caching and 50% from batch processing.3
The summary page now lists Sonnet 4.6 under legacy models, still at $3 and $15.2
The current rate card shows no price step between the three older models. Sonnet 4.5 and Sonnet 4 sit on the same $3 and $15 card as Sonnet 4.6.1
The tokenizer changed with Sonnet 5, which is why the per-text saving is closer to 13% than 33%.1
What the ranking cannot tell you
- What a buyer pays. These are list prices. Anthropic says volume discounts may be available for high-volume users, negotiated case by case.1
- What a task costs. Cost per task depends on token counts, the tokenizer, cache hit rates and retries. The rate card shows none of them.
- Partner-cloud prices. Bedrock and Google Cloud set their own rates. This ranking covers Anthropic’s own API.
- Tax. The documentation makes no tax statement for direct API billing.
- Quality. Price ranks cost, not capability. A cheaper model that fails a task pays for the attempt and the rerun.
This ranking changes when Anthropic publishes a new Sonnet rate card. New research arrives through our email briefing.
Frequently asked questions
How much does the Claude API cost?
It depends on the model. Input rates run from $0.80 per million tokens on Haiku 3.5 to $15 on Opus 4.1, and output from $4 to $75.1
Sonnet sits in the middle at $2 input and $10 output on Sonnet 5.1
New users receive a small amount of free credits to test the API.1
Some tools add their own charges. Web search costs $10 per 1,000 searches on top of standard token costs.1
Can I use Claude Sonnet without paying for the API?
Yes, through Claude’s apps rather than the API. Anthropic’s plan table includes Sonnet on the Free plan as well as on Pro and Max.2
Pro costs $17 a month with an annual subscription, or $20 billed monthly.2
Those plans carry usage limits rather than per-token rates, so they don’t compare directly with API pricing.
Is Sonnet cheaper through AWS or Microsoft?
Not through Anthropic’s own marketplace channels. Claude Platform on AWS rates token usage at the standard per-model rates, applies any negotiated discount and bills the result as Claude Consumption Units at $0.01 each.1
Claude in Microsoft Foundry bills the same way through the Azure Marketplace.1
Bedrock and Google Cloud set their own regional prices, so check their pricing pages directly.1
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


