LLM API pricing ranked: OpenRouter September 2026
Mistral's mistral-nemo has the lowest output price of any major-lab model on OpenRouter's list, at US$0.03 per million output tokens and US$0.019 input. Rank on output, then price your own input-to-output mix: the second-ranked listing costs more than the fifth on an input-heavy workload.

The cheapest LLM API from a major lab on OpenRouter’s model list, checked 28 September 2026, is Mistral’s mistralai/mistral-nemo: US$0.03 per million output tokens, with input at US$0.019 per million.1 Second is mistralai/ministral-8b-2512:batch at US$0.075 for both.1
The ranking orders 318 paid listings from 13 labs by output price alone, so read the input column before you choose. It sits within our model infrastructure research.
TL;DR
- Mistral holds the top two places and 6 of the 22 listings priced at or under US$0.20 per million output tokens. OpenAI holds 5, Google and Alibaba Qwen 3 each, Meta 2, and Amazon, Cohere and Zhipu one each.1
- Anthropic, xAI, Moonshot AI, MiniMax and DeepSeek have nothing in that band. Anthropic’s lowest output price is US$2.50, about 83 times the top of the list.1
- Output prices run from US$0.03 to US$600.00 per million tokens across the 318 listings. The top end is
openai/o1-pro.1 - Output rank and bill rank part ways. On 10 million input and 2 million output tokens, the second-ranked listing costs US$0.90 and the fifth-ranked
openai/gpt-oss-20bcosts US$0.36.1 - Every figure is a listed price in US dollars, and the ranking adds no OpenRouter plan fee. Its pricing page lists 5.5% on Standard and 8% on Business.2
How this ranking was built
This ranking orders the major labs’ models by output price per million tokens from OpenRouter’s model list, checked 28 September 2026, as the single source at full precision.1 The source is OpenRouter’s public model API, which needs no key.1 Each listing carries two prices in US dollars per token: pricing.completion for output and pricing.prompt for input.1 Glamdring Research multiplied both by 1,000,000 and changed nothing else, so a figure such as US$0.09375 stays as listed.1
The inclusion rule keeps paid listings from 13 labs: OpenAI, Mistral, DeepSeek, Meta, Alibaba Qwen, Amazon, Cohere, Google, Zhipu, MiniMax, Moonshot AI, Anthropic and xAI. That leaves 318 rows.1 Free listings are out. OpenRouter’s pricing page counts more than 25 free models on its Free plan, capped at 50 requests a day.2 A listing with a :batch suffix is its own row with its own price, so it’s ranked on its own.1
The sort is output price, lowest first. Where two listings share an output price, the lower input price goes first, then the model id in alphabetical order. Nothing was rounded, weighted or blended.1 The research standard behind single-source rankings is published separately.
Which LLM APIs cost the least per million output tokens?
Mistral’s mistralai/mistral-nemo costs the least, at US$0.03 per million output tokens and US$0.019 per million input tokens on OpenRouter’s model list, checked 28 September 2026.1 The 22 cheapest listings by output price all sit at or under US$0.20 per million, and they come from eight labs.1
| # | Model id | Lab | Output (US$ per million tokens) | Input (US$ per million tokens) | Context (tokens) |
|---|---|---|---|---|---|
| 1 | mistralai/mistral-nemo | Mistral | 0.03 | 0.019 | 131,072 |
| 2 | mistralai/ministral-8b-2512:batch | Mistral | 0.075 | 0.075 | 262,144 |
| 3 | meta-llama/llama-3.1-8b-instruct | Meta | 0.08 | 0.05 | 131,072 |
| 4 | mistralai/mistral-small-24b-instruct-2501 | Mistral | 0.08 | 0.05 | 32,768 |
| 5 | openai/gpt-oss-20b | OpenAI | 0.09 | 0.018 | 131,072 |
| 6 | google/gemma-3-4b-it | 0.10 | 0.05 | 131,072 | |
| 7 | mistralai/ministral-3b-2512 | Mistral | 0.10 | 0.10 | 131,072 |
| 8 | openai/gpt-oss-20b:batch | OpenAI | 0.112 | 0.024 | 131,072 |
| 9 | qwen/qwen3.7-flash | Alibaba Qwen | 0.13 | 0.03 | 1,000,000 |
| 10 | openai/gpt-oss-120b:batch | OpenAI | 0.136 | 0.0296 | 131,072 |
| 11 | amazon/nova-micro-v1 | Amazon | 0.14 | 0.035 | 128,000 |
| 12 | cohere/command-r7b-12-2024 | Cohere | 0.15 | 0.0375 | 128,000 |
| 13 | google/gemma-3-12b-it | 0.15 | 0.05 | 131,072 | |
| 14 | qwen/qwen3.5-9b | Alibaba Qwen | 0.15 | 0.10 | 262,144 |
| 15 | mistralai/ministral-8b-2512 | Mistral | 0.15 | 0.15 | 262,144 |
| 16 | meta-llama/llama-guard-4-12b | Meta | 0.18 | 0.18 | 163,840 |
| 17 | openai/gpt-5-nano:batch | OpenAI | 0.20 | 0.025 | 400,000 |
| 18 | google/gemini-2.5-flash-lite:batch | 0.20 | 0.05 | 1,048,576 | |
| 19 | openai/gpt-4.1-nano:batch | OpenAI | 0.20 | 0.05 | 1,047,576 |
| 20 | z-ai/glm-5.3-flash:batch | Zhipu | 0.20 | 0.06 | 1,048,576 |
| 21 | qwen/qwen-2.5-7b-instruct | Alibaba Qwen | 0.20 | 0.10 | 32,768 |
| 22 | mistralai/ministral-14b-2512 | Mistral | 0.20 | 0.20 | 262,144 |
Every row comes from the same capture and the same two fields.1 The 23rd listing is meta-llama/llama-3.2-1b-instruct at US$0.201, so the table stops at a clean US$0.20 line.1
Context windows in the top 22 run from 32,768 tokens to 1,048,576. qwen/qwen3.7-flash, ninth at US$0.13, is the cheapest listing by output price with a context window of a million tokens or more.1 A reader who needs long documents in one request can start there.
Seven of the 22 are :batch listings, and a batch row isn’t always the cheaper twin.1 mistralai/ministral-8b-2512:batch charges US$0.075 against US$0.15 for the standard listing, and openai/gpt-5-nano:batch charges US$0.20 output against US$0.40. Both are half.1 openai/gpt-oss-20b:batch runs the other way. Its US$0.112 output and US$0.024 input both sit above the standard listing’s US$0.09 and US$0.018.1 deepseek/deepseek-v4.1-flash:batch is dearer than its standard listing too, at US$0.336 output against US$0.29.1 The captured list doesn’t say what the suffix changes in delivery, so treat each batch row as a separate product with a separate price.
What is the cheapest model from each lab?
By output price, Mistral’s cheapest listing is the cheapest of any lab at US$0.03 per million tokens and Anthropic’s is the dearest at US$2.50, on OpenRouter’s model list checked 28 September 2026.1 Four labs have nothing below US$1.00 per million output tokens: MiniMax, xAI, Moonshot AI and Anthropic.1
| Lab | Cheapest listing by output price | Output (US$ per million tokens) | Input (US$ per million tokens) |
|---|---|---|---|
| Mistral | mistralai/mistral-nemo | 0.03 | 0.019 |
| Meta | meta-llama/llama-3.1-8b-instruct | 0.08 | 0.05 |
| OpenAI | openai/gpt-oss-20b | 0.09 | 0.018 |
google/gemma-3-4b-it | 0.10 | 0.05 | |
| Alibaba Qwen | qwen/qwen3.7-flash | 0.13 | 0.03 |
| Amazon | amazon/nova-micro-v1 | 0.14 | 0.035 |
| Cohere | cohere/command-r7b-12-2024 | 0.15 | 0.0375 |
| Zhipu | z-ai/glm-5.3-flash:batch | 0.20 | 0.06 |
| DeepSeek | deepseek/deepseek-v4-flash | 0.28 | 0.14 |
| MiniMax | minimax/minimax-m2.5 | 1.08 | 0.27 |
| xAI | x-ai/grok-4.3:batch | 2.00 | 1.00 |
| Moonshot AI | moonshotai/kimi-k2.5 | 2.25 | 0.45 |
| Anthropic | anthropic/claude-haiku-4.5:batch | 2.50 | 0.50 |
The lab order shifts when the field changes. DeepSeek’s deepseek/deepseek-v4-flash-0731 has the third-lowest input price of all 318 listings at US$0.021, but its output price is US$0.32.1 DeepSeek’s lowest output price belongs to deepseek/deepseek-v4-flash, at US$0.28 with US$0.14 input.1 Meta flips the same way. meta-llama/llama-3.2-1b-instruct has Meta’s lowest input price at US$0.027, while meta-llama/llama-3.1-8b-instruct has its lowest output price at US$0.08.1
xAI’s two cheapest listings tie on both fields: x-ai/grok-4.3:batch and x-ai/grok-build-0.1, each at US$1.00 input and US$2.00 output.1 MiniMax’s cheapest output price, US$1.08 on minimax/minimax-m2.5, sits just under the US$1.10 of minimax/minimax-01, the MiniMax listing with the lowest input price.1
How do you read a per-token price?
A per-token price is what one token costs. OpenRouter’s model list states it in US dollars per token, separately for input, the tokens you send, and output, the tokens the model returns.1,3 Multiply by 1,000,000 to get the per-million figure used here. Then price a workload as input tokens times the input price, plus output tokens times the output price.
For mistralai/mistral-nemo, US$0.03 per million output tokens is US$0.00000003 per token.1 TypingMind’s pricing explainer puts a token at roughly three-quarters of an English word and counts prompts, uploaded documents and images as input.3
The table below prices one assumed workload: 10 million input tokens and 2 million output tokens. The five-to-one mix is an illustration chosen for this example, not a measured average. Every price comes from the same capture.1
| Listing | Output rank | Input cost (US$) | Output cost (US$) | Total (US$) |
|---|---|---|---|---|
mistralai/mistral-nemo | 1 | 0.19 | 0.06 | 0.25 |
mistralai/ministral-8b-2512:batch | 2 | 0.75 | 0.15 | 0.90 |
meta-llama/llama-3.1-8b-instruct | 3 | 0.50 | 0.16 | 0.66 |
openai/gpt-oss-20b | 5 | 0.18 | 0.18 | 0.36 |
qwen/qwen3.7-flash | 9 | 0.30 | 0.26 | 0.56 |
anthropic/claude-sonnet-4.6 | below top 22 | 30.00 | 30.00 | 60.00 |
openai/o1-pro | 318 | 1,500.00 | 1,200.00 | 2,700.00 |
mistralai/mistral-nemo stays cheapest on that mix. openai/gpt-oss-20b has the lower input price, US$0.018 against US$0.019, so it overtakes only when input tokens exceed 60 times output tokens.1 The second-ranked listing falls behind four others because its US$0.075 input price is higher than any of theirs.1 Take your ratio from your own request logs before relying on either order.
The per-million figure starts the estimate, and the invoice can differ. TypingMind’s own estimator says prompt caching, batched calls, volume discounts, reasoning-token overhead and provider billing rules can move the final cost.3 Our explainer on what controls AI inference cost covers the cost drivers that sit underneath these prices.
How far apart are input and output prices?
Output costs more than input on most listings, and the multiple runs from below one to more than 16 on OpenRouter’s model list, checked 28 September 2026.1 A low input price buys little when a model writes long answers, so the multiple tells you which column governs the bill.
Every Anthropic listing in the capture prices output at five times input. anthropic/claude-haiku-4.5 is US$1.00 and US$5.00, anthropic/claude-sonnet-4.6 is US$3.00 and US$15.00, and anthropic/claude-opus-5 is US$5.00 and US$25.00.1 openai/gpt-5 and google/gemini-2.5-pro both list US$1.25 input and US$10.00 output, eight to one. google/gemini-3.1-pro-preview is six to one at US$2.00 and US$12.00.1
At the flat end, Mistral’s 2512 Ministral listings charge the same for both directions: US$0.10 for mistralai/ministral-3b-2512, US$0.15 for mistralai/ministral-8b-2512 and US$0.20 for mistralai/ministral-14b-2512. meta-llama/llama-3.1-70b-instruct is flat at US$0.40.1 openai/gpt-5-image-mini inverts the pattern, at US$2.50 input and US$2.00 output.1 The widest gap among the listings named here is deepseek/deepseek-v4-pro-0813, at US$0.2523 input and US$4.20 output, more than 16 to one.1
TypingMind’s explanation for the usual gap is compute. It says each generated token needs its own forward pass, while input tokens are processed in parallel.3 That is a general account, and OpenRouter’s list doesn’t give a reason for any single price.
Why do other published LLM prices disagree?
Published LLM prices disagree for three reasons: they come from different hosts, they were checked on different dates, and some quote input price alone. Where host, date and model match, the figures can agree exactly. Holori lists Claude Sonnet 4.6 on AWS Bedrock at $3.00 input and $15.00 output per million tokens, the same as OpenRouter’s anthropic/claude-sonnet-4.6.4,1
Host changes the number. TypingMind lists GLM-5.2 through Alibaba (China) at $1.10 input and $3.85 output, while OpenRouter lists z-ai/glm-5.2 at US$0.6496 and US$2.0416.3,1 The same TypingMind page puts Kimi K3 at $2.83 and $14.13, against US$3.00 and US$15.00 for moonshotai/kimi-k3 on OpenRouter.3,1 The gap is small for Qwen3.7 Flash: $0.0296 and $0.1185 at TypingMind, US$0.03 and US$0.13 on OpenRouter.3,1
Date and naming change it again. Guzli’s January 2026 table puts DeepSeek V3 at $0.27 input and $1.10 output.5 None of the DeepSeek listings in the September capture is called simply V3. The capture carries deepseek/deepseek-chat at US$0.2574 and US$1.0287, deepseek/deepseek-chat-v3-0324 at US$0.29 and US$1.14, and deepseek/deepseek-chat-v3.1 at US$0.25 and US$0.95.1 A reader matching that name to a price has three candidates and documents from January and September.
The field matters most. llm-stats names Nova 2 Lite the most affordable Bedrock model at $0.300 per million input tokens.6 OpenRouter lists amazon/nova-2-lite-v1 at the same US$0.30 input, with output at US$2.50, a figure a ranking on input never shows.1 Compare host, checked date and field before two prices go side by side.
What the ranking cannot tell you
OpenRouter lists mistralai/mistral-nemo at US$0.03 per million output tokens. That is a listed price, not a measured bill.1 The ranking says nothing about output quality, speed, uptime or rate limits. OpenRouter’s pricing page describes rate limits on its Standard plan as passthrough from the providers.2
Fees sit outside the ranked field. OpenRouter’s pricing page lists a 5.5% platform fee on Standard, 8% on Business and discounts on Enterprise. Bring-your-own-key use on Standard and Business carries no fee on the first US$25,000 of list-price inference a month and 5% after.2 That page also poses its own question about whether VAT or GST is included, but the answer wasn’t in the capture, so the tax status of these prices is unconfirmed.2 Plan fees are covered in our OpenRouter pricing research.
Two more limits matter when the numbers travel. The captured fields hold one input price and one output price per listing, so cached-input rates don’t appear. Holori lists a separate cached-input price of $0.300 per million for Claude Sonnet 4.6 on Bedrock.4 And these are OpenRouter’s prices. The labs’ own price pages weren’t part of this check, so confirm a direct contract against the lab’s current quote.
This is Glamdring Research’s first capture of the list, so nothing here shows a price moving. The ranking changes when OpenRouter’s list does. The next capture will carry a change note and its sources: subscribe to the email briefing.
Frequently asked questions
Which LLM has the cheapest API?
On OpenRouter’s model list checked 28 September 2026, Mistral’s mistralai/mistral-nemo has the lowest output price of any major-lab listing, US$0.03 per million tokens, with input at US$0.019.1 On input price alone, openai/gpt-oss-20b is lowest at US$0.018, with output at US$0.09.1 Which is cheaper for a given job depends on its mix. openai/gpt-oss-20b costs less only when input tokens exceed 60 times output tokens.
How much is 1 million tokens in LLM?
Across 318 paid major-lab listings on OpenRouter, checked 28 September 2026, a million output tokens cost from US$0.03 to US$600.00 and a million input tokens from US$0.018 to US$150.00.1 Both ends of the input range and the top of the output range belong to OpenAI: openai/gpt-oss-20b at the bottom of input and openai/o1-pro at the top of both.1 On TypingMind’s rule of three-quarters of a word per token, a million tokens is roughly 750,000 English words.3
How much does LLM cost?
An LLM API bill is input tokens times the input price plus output tokens times the output price, plus any platform fee. On an assumed 10 million input and 2 million output tokens, the 318 listings checked on 28 September 2026 run from US$0.25 for mistralai/mistral-nemo to US$2,700.00 for openai/o1-pro.1 OpenRouter’s pricing page lists a 5.5% platform fee on its Standard plan.2
How much does ChatGPT API cost?
None of the 318 listings checked on 28 September 2026 is named ChatGPT. The OpenAI listings run from openai/gpt-oss-20b at US$0.018 input and US$0.09 output to openai/o1-pro at US$150.00 and US$600.00 per million tokens.1 The listing openai/gpt-chat-latest is US$5.00 input and US$30.00 output.1 These are OpenRouter’s prices. OpenAI’s own price page wasn’t part of this check, so confirm a direct contract against it.
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


