Gemini 2.5 Pro API pricing: Google, September 2026
Gemini 2.5 Pro lists $1.25 input and $10 output per million tokens for prompts up to 200,000 tokens. Gemini 3.1 Pro Preview lists $2 and $12. Longer prompts increase both rates, on Google's prices checked 28 September 2026.

Gemini 2.5 Pro costs $1.25 per million input tokens and $10.00 per million output tokens on Google’s Gemini Developer API, for prompts up to 200,000 tokens. Above that, the rates rise to $2.50 and $15.00.1 That makes it the cheaper of Google’s two Pro text models: Gemini 3.1 Pro Preview lists $2.00 and $12.00.1 Prices are in US dollars, from Google’s Gemini API pricing page, checked 28 September 2026.
Output includes thinking tokens, so a long reasoning chain bills at the output rate.1
TL;DR
- Gemini 2.5 Pro is the cheapest Pro text model. Gemini 3.1 Pro Preview costs 60% more per input token and 20% more per output token.
- The over-200,000-token tier doubles input and lifts output by half on both models. Keeping prompts at or under 200,000 tokens keeps the input rate at half the long-prompt rate.
- Context caching cuts cached reads by 90%, but storage runs $4.50 per million tokens per hour. A Gemini 2.5 Pro cache pays for itself at four reads an hour; a Gemini 3.1 Pro Preview cache at 2.5.
- Only Gemini 2.5 Pro has a free tier. Free-tier content is used to improve Google’s products, and caching is paid-only.
- Batch and Flex halve token prices. Priority costs 1.8 times Standard. Cache reads don’t get the Batch discount, so caching saves less on batch jobs.
How this ranking was built
The source is Google’s Gemini Developer API pricing page, which Google marks as last updated 24 September 2026.1 We captured and checked it on 28 September 2026, under our published research standard.
- Ranked field: the Input price row, Paid Tier, Standard, “prompts <= 200k tokens”. Unit: USD per 1M tokens. The page doesn’t state tax treatment.
- Inclusion rule: every model section on the page with “Pro” in its name. That gives four sections: Gemini 3.1 Pro Preview, Gemini 3 Pro Image (Nano Banana Pro), Gemini 2.5 Pro and Gemini 2.5 Pro Preview TTS.1
- Merge:
gemini-3.1-pro-preview-customtoolsshares the Gemini 3.1 Pro Preview table, so it counts as one row.1 Four rows remain. - Exclusions: Gemini 3 Pro Image and Gemini 2.5 Pro Preview TTS produce images and audio, priced on different output units. They’re covered in their own section below. Two rows are ranked.
- Also outside the ranking: Gemini 2.5 Computer Use Preview carries the same token rates as Gemini 2.5 Pro but isn’t named Pro on the Developer API page.1
- Tier labels: each Pro section shows Standard, Batch, Flex and Priority tabs above four tables. We read the tables in that order. Google’s Agent Platform pricing page labels the same figures Standard, Priority and Flex/Batch, which confirms the reading.2
Which Gemini Pro model is cheapest per million tokens?
Gemini 2.5 Pro, at $1.25 per million input tokens against $2.00 for Gemini 3.1 Pro Preview.1
| Rank | Model | Status | Input, prompts ≤200k (USD per 1M tokens) | Output incl. thinking, prompts ≤200k (USD per 1M) | Input, prompts >200k (USD per 1M) | Output, prompts >200k (USD per 1M) | Free tier |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 2.5 Pro (gemini-2.5-pro) | No preview label | $1.25 | $10.00 | $2.50 | $15.00 | Input and output free of charge |
| 2 | Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) | Preview | $2.00 | $12.00 | $4.00 | $18.00 | Not available |
The gap is uneven. Gemini 3.1 Pro Preview charges 60% more on input and 20% more on output, so input-heavy work widens the difference and output-heavy work narrows it.
Take a request with 150,000 input tokens and 2,000 output tokens. On Gemini 2.5 Pro it costs $0.2075: $0.1875 of input and $0.02 of output. On Gemini 3.1 Pro Preview it costs $0.324: $0.30 of input and $0.024 of output.
Google describes Gemini 3.1 Pro Preview as “Our 3rd generation Pro model, built for multimodal understanding, agentic capabilities, and vibe-coding”, and Gemini 2.5 Pro as “A Pro model which excels at coding and complex reasoning tasks”.1 Those are Google’s claims. The price table can’t tell you whether the newer model earns its premium on your workload. Run one real workflow through both and compare cost per completed task.
Google’s Flash models aren’t ranked here. Some list lower rates than either Pro model: Gemini 3.8 Flash lists $0.75 per million input tokens through December 31, 2026.1 Our full Gemini API price list covers them.
What does Gemini 2.5 Pro cost above 200,000 tokens?
Once a prompt passes 200,000 tokens, Gemini 2.5 Pro costs $2.50 per million input tokens and $15.00 per million output tokens.1 Gemini 3.1 Pro Preview moves to $4.00 and $18.00.1 Cache reads double too: $0.125 to $0.25 on Gemini 2.5 Pro, and $0.20 to $0.40 on Gemini 3.1 Pro Preview.1
The trigger is prompt length. Google’s table conditions the output price on “prompts > 200k”, not on output length.1 A long prompt makes every output token dearer, too.
At 300,000 input tokens and 2,000 output tokens, Gemini 2.5 Pro costs $0.78: $0.75 of input and $0.03 of output. Gemini 3.1 Pro Preview costs $1.236: $1.20 of input and $0.036 of output. The estimate assumes the whole prompt bills at the higher rate, which is how Google’s table reads. Confirm this against your first invoice.
The constraint is the 200,000-token line. Retrieval that trims a 300,000-token prompt to 150,000 tokens cuts the Gemini 2.5 Pro input charge from $0.75 to $0.1875, because it halves the tokens and the rate. Our explainer on what controls AI inference cost covers the wider set of levers.
How much does context caching save on Gemini Pro?
On Glamdring Research’s calculation, a Gemini 2.5 Pro cache pays for itself at four reads an hour on prompts up to 200,000 tokens, and a Gemini 3.1 Pro Preview cache at 2.5 reads an hour. Cached reads cost a tenth of the input rate, but storage costs $4.50 per million tokens per hour on both models.1
| Model | Cache read, prompts ≤200k (USD per 1M) | Cache read, prompts >200k (USD per 1M) | Storage (USD per 1M tokens per hour) | Break-even reads per hour, ≤200k | Break-even reads per hour, >200k |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | $0.125 | $0.25 | $4.50 | 4.0 | 2.0 |
| Gemini 3.1 Pro Preview | $0.20 | $0.40 | $4.50 | 2.5 | 1.25 |
The break-even is storage price divided by the saving per read: $4.50 ÷ ($1.25 − $0.125) = 4.0 for Gemini 2.5 Pro. The cache size cancels out, so the figure holds for any cache. It assumes storage bills pro rata by the hour and that each cached read replaces a full-price read of the same tokens.
A worked case: a 150,000-token document cached on Gemini 2.5 Pro costs $0.675 an hour to store. Each read costs $0.01875 instead of $0.1875. At ten reads an hour, uncached input costs $1.875 and cached input costs $0.8625, a 54% saving. At two reads an hour, the cache costs more than it saves. The estimate counts storage and reads only.
Caching is a paid-tier feature. Google lists “Access to Context caching” under Paid, and the Gemini 2.5 Pro free tier shows caching as not available.1
Can you use the Gemini Pro API for free?
Yes for Gemini 2.5 Pro, no for Gemini 3.1 Pro Preview. Google’s Gemini 2.5 Pro table lists input and output as “Free of charge” on the free tier.1 The Gemini 3.1 Pro Preview table lists them as “Not available”.1
The free tier has three costs that aren’t in dollars:
- Google uses free-tier content to improve its products. The free tier says “Content used to improve our products”; the paid tier says it isn’t.1
- No caching or grounding on Gemini 2.5 Pro. Both show “Not available” on the free tier.1 Gemini 3.1 Pro Preview grounding “Can be tested in Google AI Studio”.1
- Lower, unpublished limits. Google describes the free tier as “Limited access to certain models”1 and points users to Google AI Studio for their active rate limits.3 The captured pages give no free-tier request numbers for Gemini 2.5 Pro.
Google AI Studio usage is free of charge in all available regions.1 Moving to paid means linking a billing account, which puts the project on Tier 1 with a $250 billing tier cap. Depending on billing history, Tier 1 also carries a spend rate limit of $10 per 10 minutes.3
What do Batch, Flex and Priority cost on Pro models?
Batch and Flex cost half the Standard token price on both Pro text models, and Priority costs 1.8 times Standard.1
| Model | Tier | Input ≤200k | Output ≤200k | Input >200k | Output >200k | Cache read ≤200k | Cache storage per hour |
|---|---|---|---|---|---|---|---|
| Gemini 2.5 Pro | Standard | $1.25 | $10.00 | $2.50 | $15.00 | $0.125 | $4.50 |
| Gemini 2.5 Pro | Batch or Flex | $0.625 | $5.00 | $1.25 | $7.50 | $0.125 | $4.50 |
| Gemini 2.5 Pro | Priority | $2.25 | $18.00 | $4.50 | $27.00 | $0.225 | $8.10 |
| Gemini 3.1 Pro Preview | Standard | $2.00 | $12.00 | $4.00 | $18.00 | $0.20 | $4.50 |
| Gemini 3.1 Pro Preview | Batch or Flex | $1.00 | $6.00 | $2.00 | $9.00 | $0.20 | $4.50 |
| Gemini 3.1 Pro Preview | Priority | $3.60 | $21.60 | $7.20 | $32.40 | $0.36 | $8.10 |
All figures USD per 1M tokens; storage per 1M tokens per hour.1
Cache reads keep their Standard price under Batch and Flex; Google marks the Gemini 3.1 Pro Preview figure “Same as Standard”.1 That changes the caching arithmetic. On Batch, a Gemini 2.5 Pro cache needs nine reads an hour to break even ($4.50 ÷ ($0.625 − $0.125)), and a Gemini 3.1 Pro Preview cache needs 5.625.
Batch jobs have their own limits: 100 concurrent batch requests, and 5,000,000 enqueued tokens per Pro model at Tier 1, rising to 500,000,000 at Tier 2 and 1,000,000,000 at Tier 3.3 Priority’s default rate limits are 0.3 times the standard limit.3
What do the Pro image and speech models cost?
Gemini 3 Pro Image charges $2.00 per million input tokens and $12.00 per million text-and-thinking output tokens, the same as Gemini 3.1 Pro, plus $120.00 per million image output tokens.1 Google converts that to $0.134 per 1K or 2K image and $0.24 per 4K image. An input image counts as 560 tokens, or $0.0011.1 On Batch, text drops to $1.00 in and $6.00 out, and images to $0.067 per 1K or 2K image and $0.12 per 4K image.1 The table lists no free tier and no over-200,000-token tier.
Gemini 2.5 Pro Preview TTS charges $1.00 per million text input tokens and $20.00 per million audio output tokens, or $0.50 and $10.00 on Batch.1 It has Standard and Batch tabs only, and no free tier.1
Why do other published Gemini Pro prices disagree?
Most disagreement comes from mixing price lists, tiers or prompt lengths. Four patterns explain it.
- Two Google price lists. Google’s Developer API page warns that its prices “may differ” from those on Gemini Enterprise Agent Platform.1 On 28 September 2026 the Pro token rates matched.2 Grounding allowances did not: the Developer API lists 1,500 free Google Search requests per day for Gemini 2.5 Pro, then $35 per 1,000 grounded prompts1, while Agent Platform says Gemini 2.5 Pro “includes 10,000 Grounding Prompts per day at no additional charge”.2
- Tab confusion. A Batch figure ($0.625) or a Priority figure ($2.25) quoted as “the” Gemini 2.5 Pro input price is off by half or by 80%.1
- Ranges. Search results for this query quote Gemini 2.5 Pro as a range, such as “$1.25-$2.50 per 1M input tokens”. Both ends are right; they’re the two prompt-length tiers.
- Schedules on other models. Gemini 3.8 Flash lists introductory prices “through December 31, 2026” and higher prices from 1 January 2027.1 The Pro tables carry no dated change.
What the ranking cannot tell you
- Model quality. The ranking sorts by list price. Google’s model descriptions are company claims.
- Your bill. Thinking tokens bill as output, and their count varies by request.1 A cheaper rate can still produce a larger bill.
- Tax and currency. Prices are USD. The Developer API page doesn’t state tax treatment. Agent Platform says non-USD payers see prices in their currency on Cloud Platform SKUs.2
- Change since the last release. We hold no earlier dated capture of this page, so we can’t report what moved before 24 September 2026. This ranking refreshes when Google next updates the page.
- Preview terms. Gemini 3.1 Pro Preview is a preview model, and Google applies tighter rate limits to preview models.3 Google also states that rate limits “are not guaranteed”.3
More model infrastructure research sits in the same field. New price checks go out in the free research briefing.
Frequently asked questions
Do thinking tokens count as output tokens on Gemini Pro?
Yes. Google’s Gemini 2.5 Pro and Gemini 3.1 Pro Preview tables label the output row “Output price (including thinking tokens)”.1 A response with 500 visible tokens and 4,500 thinking tokens bills as 5,000 output tokens: $0.05 on Gemini 2.5 Pro at the short-prompt rate.
Does Google Search grounding cost extra on Gemini Pro?
Yes, after a free allowance on the paid tier. Gemini 2.5 Pro includes 1,500 requests per day, then costs $35 per 1,000 grounded prompts.1 Gemini 3.1 Pro Preview includes 5,000 free search requests per month, shared across all Gemini 3.x models, then costs $14 per 1,000 requests.1 On Gemini 3.x, one request can trigger several Google Search queries, and each query is charged.1
How are PDFs billed on the Gemini Pro API?
Google bills PDF tokens, the DOCUMENT modality, at the image token rate.1 On Gemini 2.5 Pro and Gemini 3.1 Pro Preview, the input row carries one price per prompt-length tier, and Google’s Agent Platform lists that price as covering text, image, video and audio input.2
What does Gemini 3 Pro cost on the API?
Google’s Developer API pricing page, checked 28 September 2026, has no separate Gemini 3 Pro text-model table. The third-generation Pro text model it prices is Gemini 3.1 Pro Preview, at $2.00 per million input tokens and $12.00 per million output tokens for prompts up to 200,000 tokens.1 Gemini 3 Pro Image is the other third-generation Pro entry, and Google prices its text input and output the same as Gemini 3.1 Pro.1
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


