Glamdring

Model infrastructure

Claude API pricing per model: Anthropic, September 2026

Haiku 4.5 is the cheapest current Claude API model at $1 input and $5 output per million tokens; Fable 5.1 is the dearest current model at $10 and $50.

Glamdring Research16 min5 sources checked

Engraving of an exploded view of a liquid-cooled AI accelerator module: the chip package, high-bandwidth memory stacks, cold plate and coolant fittings

Haiku 4.5 is the cheapest current Claude model on the Anthropic API, at $1 per million input tokens and $5 per million output tokens. Sonnet 5 costs $2 and $10, Opus 5.5 costs $4 and $20, and Fable 5.1 costs $10 and $50.1 Those are Anthropic’s published rates, checked 28 September 2026, in US dollars.1 The Batch API takes 50% off both input and output.1

Output costs five times input on every model Anthropic prices, so a long answer moves the bill faster than a long prompt.

TL;DR

  • Opus 5.5, at $4 input and $20 output, costs less per token than all five Opus models Anthropic now files as legacy, which list at $5 and $25.1,2
  • A 5-minute cache write costs 1.25 times the input rate and a 1-hour write costs twice it. Cache reads cost 10% of input, except on Opus 5.5 (5%) and Fable 5.1 (2.5%).1
  • Our worked request, a 10,000-token prompt with a 1,000-token reply, costs $0.030 on Sonnet 5. With 8,000 prompt tokens cached, a repeat of it on Opus 5.5 costs $0.0296, slightly less than the uncached Sonnet 5 request.
  • Fable 5.1 matches Fable 5 on input, output and cache writes, and reads from cache at $0.25 against Fable 5’s $1.1
  • Claude subscriptions sell app and Claude Code usage under rolling limits with no published token allowance, so there’s no exact break-even against the API.2 Light, irregular or automated use favours the API. All-day interactive use favours a plan.
  • More cost and serving analysis sits in our model infrastructure coverage.

How this ranking was built

Every per-token figure comes from Anthropic’s API pricing documentation, captured at 15:05 AEST on 28 September 2026.1 Anthropic’s public pricing page, captured the same minute, shows the same input, output, 5-minute cache write and cache read rates for all 12 models it lists.2

  • Ranked field: the documentation’s Input column, in US dollars per million tokens (MTok). All prices on that page are in USD.1
  • Order: lowest input price first. Models with the same input price share a rank.
  • Inclusion: the documentation prices 16 models. Twelve also appear on the pricing page, four under “Latest models” and eight under “Legacy models”.2 Those 12 are ranked. The other six follow the table, unranked. Nothing was merged.
  • Cache columns: the documentation prints the 5-minute write, the 1-hour write and the read rate, which it calls “Hits and refreshes”, for every model.1
  • Batch columns: printed per model in the documentation’s batch table, not computed by us.1
  • Tax: Anthropic states no tax position for API token rates. It says its plan prices exclude applicable tax.2

These are list prices Anthropic publishes about its own service. They’re company statements, not measured market prices. Our research methodology sets out how we capture, date and check figures like these.

Claude API prices by model, ranked by input price

Haiku 4.5 has the lowest input price of the 12 models on Anthropic’s pricing page, at $1 per MTok. Fable 5.1 and Fable 5 share the highest, at $10.1

RankModelPricing page statusInput ($/MTok)Output ($/MTok)5-min cache write ($/MTok)1-hour cache write ($/MTok)Cache read ($/MTok)Batch input ($/MTok)Batch output ($/MTok)
1Haiku 4.5Latest151.2520.100.502.50
2Sonnet 5Latest2102.5040.2015
3=Sonnet 4.6Legacy3153.7560.301.507.50
3=Sonnet 4.5Legacy3153.7560.301.507.50
5Opus 5.5Latest420580.20210
6=Opus 5Legacy5256.25100.502.5012.50
6=Opus 4.8Legacy5256.25100.502.5012.50
6=Opus 4.7Legacy5256.25100.502.5012.50
6=Opus 4.6Legacy5256.25100.502.5012.50
6=Opus 4.5Legacy5256.25100.502.5012.50
11=Fable 5.1Latest105012.50200.25525
11=Fable 5Legacy105012.50201525

Standard, cache and batch rates come from Anthropic’s documentation, checked 28 September 2026.1

Three ratios hold on every row. Output is five times input, a 5-minute cache write is 1.25 times input, and batch rates are half the standard rates. Cache reads are where the rows split.

Within each family, the newest model is never dearer per token than the one before it. Opus 5.5 lists below Opus 5 on every price, and Sonnet 5 lists below Sonnet 4.6 on every price. Fable 5.1 ties Fable 5 on everything except the cache read, which is a quarter of the price.1

Six more models appear in the documentation’s table but not on the pricing page.1 Mythos 5.1 and Mythos 5 carry a marker the capture doesn’t label, and BenchLM lists Mythos 5 as limited access.3 Opus 4.1, Opus 4, Sonnet 4 and Haiku 3.5 appear without links to model pages.

Model (documentation only)Input ($/MTok)Output ($/MTok)5-min cache write ($/MTok)1-hour cache write ($/MTok)Cache read ($/MTok)Batch input ($/MTok)Batch output ($/MTok)
Mythos 5.1105012.50200.25525
Mythos 5105012.50201525
Opus 4.1157518.75301.507.5037.50
Opus 4157518.75301.507.5037.50
Sonnet 43153.7560.301.507.50
Haiku 3.50.80411.600.080.402

Haiku 3.5’s $0.80 input rate is the lowest anywhere in Anthropic’s table, but the pricing page doesn’t list it.1 Haiku 4.5 is the cheapest model Anthropic presents on that page.

What do cache writes, cache reads and batch processing cost?

A 5-minute cache write costs 1.25 times a model’s input rate and a 1-hour write costs 2 times it. A cache read costs 0.1 times input, except 0.05 times on Opus 5.5 and 0.025 times on Fable 5.1 and Mythos 5.1.1 The Batch API halves input and output rates.1

On Sonnet 5 that means $2.50 per MTok to write a 5-minute cache, $4 for a 1-hour cache and $0.20 to read either, against $2 for plain input.1

The economics follow from those ratios. At 10% of input, a cache hit repays a 5-minute write after one read and a 1-hour write after two reads.1 Opus 5.5 reads at $0.20 and Fable 5.1 at $0.25, so every later read on those two models saves more.1 The risk runs the other way. A block written to cache and never read inside its lifetime costs 25% more than plain input on a 5-minute cache, and 100% more on a 1-hour cache.

Anthropic’s pricing page shows 5-minute cache prices only.2 The 1-hour rates in the ranked table come from the documentation.

The discounts combine. Anthropic says the caching multipliers stack with the Batch API discount and with data residency pricing.1 Batch processing is asynchronous, so it suits work that doesn’t need an immediate answer.1 Fast mode can’t be used with the Batch API.1

What does one typical request cost on each model?

A request with a 10,000-token prompt and a 1,000-token reply costs $0.030 on Sonnet 5 at standard rates, $0.015 on Haiku 4.5, $0.060 on Opus 5.5 and $0.150 on Fable 5.1. That’s our arithmetic on Anthropic’s rates, checked 28 September 2026: input tokens times the input rate, plus output tokens times the output rate, divided by one million.1

The request shape is our illustration. Picture a fixed system prompt carrying instructions and reference material, a short question and a one-page answer. Anthropic’s rough guide is that one token is about four characters or 0.75 English words, which puts the prompt near 7,500 words.1

Here is the Sonnet 5 arithmetic in full.

  • Input: 10,000 × $2 ÷ 1,000,000 = $0.020.
  • Output: 1,000 × $10 ÷ 1,000,000 = $0.010.
  • Total: $0.030. A thousand such requests cost $30.

Now cache 8,000 of the 10,000 prompt tokens for five minutes.

  • The first request writes the cache: (8,000 × $2.50 + 2,000 × $2 + 1,000 × $10) ÷ 1,000,000 = $0.034.
  • A repeat inside five minutes reads it: (8,000 × $0.20 + 2,000 × $2 + 1,000 × $10) ÷ 1,000,000 = $0.0156.
  • Two requests cost $0.0496 with caching and $0.060 without.

Sent through the Batch API with no cache, the same Sonnet 5 request costs $0.015. Batched with a cache hit, and with the multipliers stacking as Anthropic describes, the repeat request costs $0.0078.1

The same scenarios on every ranked model:

ModelStandard request ($)First request, writes 8,000-token cache ($)Repeat request, reads cache ($)Batch request, no cache ($)
Haiku 4.50.0150.0170.00780.0075
Sonnet 50.0300.0340.01560.015
Sonnet 4.6 or 4.50.0450.0510.02340.0225
Opus 5.50.0600.0680.02960.030
Opus 5, 4.8, 4.7, 4.6 or 4.50.0750.0850.0390.0375
Fable 5.10.1500.1700.0720.075
Fable 50.1500.1700.0780.075

Opus 5.5’s cheap cache read changes the model decision. Its repeat request, at $0.0296, costs slightly less than a standard Sonnet 5 request at $0.030 and well under an uncached Sonnet 4.6 request at $0.045. Opus 5.5 and Sonnet 5 both read cached tokens at $0.20 per MTok, so once most of the prompt is cached, the gap between them comes down to the fresh input and the output.1 The decision changes when a workload reuses a long fixed prompt many times inside the cache lifetime.

Use one real workflow to test this. Log a day of actual input, cached and output tokens, then price them against the ranked table.

What else changes the per-token bill?

Three published options change the rate itself. US-only inference costs 1.1 times standard on Claude 4.6 and later models, across input, output, cache writes and cache reads.1 Fast mode costs $8 input and $40 output per MTok on Opus 5.5, and $10 and $50 on Opus 5 and Opus 4.8.1 The Batch API halves the rate.1

Context length doesn’t. Claude 4.6 and later models bill a 900k-token request at the same per-token rate as a 9k-token request.1

Applied to the worked example, US-only inference lifts Sonnet 5 to $2.20 input and $11 output per MTok, so the request costs $0.033. Fast mode doubles the Opus 5.5 request from $0.060 to $0.120.

Tools bill on top of tokens.

  • Web search costs $10 per 1,000 searches, plus standard token costs for the content it brings back.1
  • Web fetch adds no charge beyond the tokens the fetched content adds to the conversation.1
  • Code execution is free when the request also includes web search or web fetch.1 Otherwise the documentation gives each organisation 1,550 free hours a month and bills $0.05 per hour per container after that.1 The pricing page states the allowance as 50 free hours a day.2
  • Managed Agents sessions cost $0.08 per session-hour of running time, plus tokens.1

Is the API or a Claude subscription cheaper?

Glamdring Research’s verdict: the API costs less for light, irregular or automated use, and a subscription suits one person who works in Claude’s apps or Claude Code for most of the day. Anthropic publishes no token allowance for its plans. Usage resets on a rolling five-hour window, paid plans add weekly limits, and chat and Claude Code draw from the same pool.2

The plans on Anthropic’s pricing page, checked 28 September 2026:

  • Free costs $0 and covers chat on web, desktop and mobile.2
  • Pro costs $17 a month billed annually ($200 up front), or $20 billed monthly.2
  • Max starts at $100 a month2 and offers 5 or 20 times Pro’s usage.2
  • A Team Standard seat costs $20 a month billed annually or $25 monthly. A Premium seat costs $100 or $125.2
  • Enterprise costs US$20 per seat a month, billed annually, plus usage at API rates.2

Plan prices exclude applicable tax.2

At API rates, Pro’s $20 monthly price buys about 667 of the example Sonnet 5 requests, or about 333 on Opus 5.5. Max’s $100 buys about 3,333 Sonnet 5 requests. Someone below those volumes pays less on the API. The arithmetic can’t say how many requests a plan’s limits allow, because Anthropic doesn’t publish that figure. Jet Admin’s Claude Code pricing guide reaches the same shape of conclusion: heavy daily use is usually cheaper on Max, and low or irregular use often costs less through the API.4

Claude Code is included in every paid plan. For heavy coding sessions, Anthropic points users to pay-as-you-go API credits through a Console account.2 Enterprise combines both models, a seat fee plus usage at API rates, so the ranked table prices the usage half of an Enterprise bill.2

Why do other published Claude prices disagree?

Most disagreement comes from model turnover. MetaCTO’s API pricing guide, dated 12 January 2026 with a May 2026 update, lists Opus 4.8 as the newest model at $5 input and $25 output.5 Anthropic’s pricing page now files Opus 4.8 under legacy models and lists Opus 5.5 at $4 and $20.2

Rules of thumb drift too. Jet Admin’s guide, which says it was re-checked against Anthropic’s page in September 2026, states that cache reads cost 10% of the standard input price.4 That holds for 10 of the 12 ranked models, but Opus 5.5 reads at 5% of input and Fable 5.1 at 2.5%.1

Plan prices differ the same way. Jet Admin lists Max 20x at $200 a month.4 Anthropic’s captured plan card prints Max as “From $100” and doesn’t show the 20x price.2

Token counts move the comparison even where rates look settled. Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text, while Sonnet 4.6 and earlier use the previous one.1 A per-token rate compared across generations therefore understates the newer model’s cost for the same text. Lift Sonnet 5’s $2 input rate by 30% and it’s about $2.60 for text that would cost $3 on Sonnet 4.6, so Sonnet 5 stays cheaper. Lift Opus 5.5’s $4 by 30% and it’s about $5.20, slightly above the $5 of Opus 4.6 and Opus 4.5, which predate the new tokenizer. The estimate assumes the 30% figure holds for your text, and Anthropic says the exact increase depends on content and workload.

What the ranking cannot tell you

This ranking orders list prices. It says nothing about capability, speed or output quality, and a lower rate doesn’t make a model cheaper per finished task. A model that needs two attempts at $2 input costs more than one that succeeds first time at $3.

  • Negotiated rates. Volume discounts are negotiated case by case,1 and sales-assisted Enterprise lists tiered incentives on committed spend.2 Neither is published.
  • Other clouds. Amazon Bedrock and Google Cloud publish their own Claude prices, which sit outside this ranking.1
  • Tool overhead. Tool definitions, tool results and search results add input tokens that a per-token table doesn’t show.
  • Change over time. This edition rests on captures from one day. It will be refreshed when Anthropic changes its pricing or adds a model.

For the serving economics underneath these list prices, see our analysis of what controls AI inference cost. Other published work is in the Glamdring research archive.

Frequently asked questions

Is the Claude API cheaper than the ChatGPT API?

On list price at the flagship tier, yes by BenchLM’s figures. BenchLM’s pricing page, captured 28 September 2026, lists GPT-5.6 Sol at $5 input and $30 output per MTok against Opus 5.5 at $4 and $20, and GPT-5.6 Terra at $2 and $12 against Sonnet 5 at $2 and $10.3 We haven’t checked the OpenAI rates against OpenAI’s own price page. Tokenizers also differ between providers, so compare cost on token counts measured from your own prompts.

Can I get the Claude API for free?

Only for testing. New API users receive a small amount of free credits.1 The free Claude plan is a separate consumer product and doesn’t provide general API access.3 Universities on Anthropic’s Education plan get dedicated API credits for student learning.2

How is the Claude API billed?

By actual monthly usage, in US dollars. Credit card and invoicing are both available, and usage is tracked in the Claude Console.1 Each request is charged for its input, output, cache write and cache read tokens at the model’s rates,1 plus tool charges such as web search.1

Why is the Claude API expensive for some workloads?

Output drives most of the bill. Every model charges five times as much for output as for input, from Haiku 4.5’s $1 and $5 to Fable 5.1’s $10 and $50.1 A task that writes long answers costs more than one that reads long documents and answers briefly. Model choice multiplies that: the example request costs ten times as much on Fable 5.1 as on Haiku 4.5. Caching repeated context and batching work that can wait are the two published ways to cut the rate, and they can be combined.1

Sources checked

  1. AnthropicChecked September 28, 2026
  2. AnthropicChecked September 28, 2026
  3. BenchLMChecked September 28, 2026
  4. Jet AdminChecked September 28, 2026
  5. MetaCTOChecked September 28, 2026

Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.