Glamdring

Model infrastructure

Gemini API pricing for every model: Google, September 2026

Gemini 2.5 Flash-Lite is the cheapest Gemini API language model at $0.10 input and $0.40 output per million tokens; Gemini 3.1 Pro Preview is the dearest at $2.00 and $12.00, rising to $4.00 and $18.00 for prompts over 200,000 tokens.

Glamdring Research23 min4 sources checked

Engraving of a hyperscale data centre hall with long aisles of server racks and overhead cable trays

Gemini 2.5 Flash-Lite is the cheapest language model on the Gemini API, at $0.10 per million input tokens and $0.40 per million output tokens on Google’s published Gemini API pricing, checked 28 September 2026.1 Gemini 3.1 Pro Preview is the most expensive, at $2.00 input and $12.00 output. Gemini 3.1 Pro Preview’s rates rise to $4.00 and $18.00 for prompts over 200,000 tokens.1 Gemini 3.8 Flash costs $0.75 and $3.75 until 31 December 2026, then doubles.1

TL;DR

  • Fourteen Gemini language models carry a per-token price. Standard paid-tier input runs from $0.10 per million tokens on Gemini 2.5 Flash-Lite to $2.00 on Gemini 3.1 Pro Preview. Output runs from $0.40 to $12.00.1
  • In our 300,000-request worked example, Gemini 3.8 Flash costs $1,237.50 a month at Standard rates now and $2,475.00 from January. Gemini 2.5 Flash-Lite runs the same load for $150.00.1 The list rate is one input to the bill; our explainer on what controls AI inference cost covers the others.
  • Three models price by prompt length: Gemini 3.1 Pro Preview, Gemini 2.5 Pro and Gemini 2.5 Computer Use Preview. Past 200,000 tokens their input rate doubles and their output rate rises by half.1 Google’s Agent Platform page says the long-context rate then applies to every token in the request, so one token over the line nearly doubles the bill.2
  • Gemini 3.8, 3.7 and 3.6 Flash sit on introductory pricing of $0.75 and $3.75 through 31 December 2026. From 1 January 2027 they cost $1.50 and $7.50.1 Until then the newest Flash costs half the input rate of the older Gemini 3.5 Flash.1
  • The free tier covers most Flash, Flash-Lite and 2.5 Pro models. The ranking excludes Gemini 3.1 Pro Preview and every image and video model, and Google may use free-tier content to improve its products.1
  • On Google’s rate tables, Batch and Flex cost half of Standard and Priority costs 1.8 times Standard.1

How this ranking was built

Glamdring Research ranked every Gemini language model on Google’s Gemini Developer API pricing page by one field: the paid-tier Input price, in US dollars per million tokens, at Standard rates.1 Where a model prices by prompt length, the rank uses the rate for prompts of 200,000 tokens or fewer. Where a model prices text and audio separately, it uses the text rate. Where a price changes on 1 January 2027, it uses the rate in force on the checked date.1 Equal input prices share a rank, and the order inside a tie follows the page.

The page was captured on 28 September 2026 and shows a last-updated date of 24 September 2026.1 It holds 35 model pricing sections. Fourteen price a language model per token with both an input and an output rate, and those fourteen are ranked. Seventeen audio, speech, image, video and embedding sections sit in a separate table below. Veo 3.1 and the two Lyria sections price per second or per song, and Gemma 4 has no paid tier.1 Prices are in US dollars. The page doesn’t state whether tax is included.1

Each model’s rates sit under tabs labelled Standard, Batch, Flex and Priority. The captured page shows the tables in that order without their labels.1 Google’s Agent Platform pricing page lists the same Priority and Flex/Batch rates for Gemini 3.1 Pro Preview and Gemini 3.8 Flash, which confirms the order.2 Ratios and workload costs in this article are Glamdring’s arithmetic on the listed rates. How we treat company-published figures is set out in our research methodology.

What does each Gemini language model cost per million tokens?

Gemini 2.5 Flash-Lite is the cheapest Gemini language model at $0.10 input and $0.40 output per million tokens. Gemini 3.1 Pro Preview is the dearest at $2.00 and $12.00. Both figures are Standard paid-tier rates from Google’s Gemini API pricing page, checked 28 September 2026.1

RankModelModel IDInput (USD per 1M tokens)Output incl. thinking (USD per 1M tokens)Cached input (USD per 1M tokens)Free tier
1Gemini 2.5 Flash-Litegemini-2.5-flash-lite$0.10$0.40$0.01Yes
2Gemini 3.1 Flash-Litegemini-3.1-flash-lite$0.25$1.50$0.025Yes
3=Gemini 3.5 Flash-Litegemini-3.5-flash-lite$0.30$2.50$0.03Yes
3=Gemini 2.5 Flashgemini-2.5-flash$0.30$2.50$0.03Yes
5Gemini 3 Flash Previewgemini-3-flash-preview$0.50$3.00$0.05Yes
6=Gemini 3.8 Flashgemini-3.8-flash$0.75*$3.75*$0.075*Yes
6=Gemini 3.7 Flashgemini-3.7-flash$0.75*$3.75*$0.075*Yes
6=Gemini 3.6 Flashgemini-3.6-flash$0.75*$3.75*$0.075*Yes
9=Gemini Robotics ER 2 Previewgemini-robotics-er-2-preview$1.00*$5.00*$0.10*Yes
9=Gemini Robotics ER 2 Streaming Previewgemini-robotics-er-2-streaming-preview$1.00*$5.00*Not listedYes
11=Gemini 2.5 Progemini-2.5-pro$1.25$10.00$0.125Yes
11=Gemini 2.5 Computer Use Previewgemini-2.5-computer-use-preview-10-2025$1.25$10.00Not listedNo
13Gemini 3.5 Flashgemini-3.5-flash$1.50$9.00$0.15Yes
14Gemini 3.1 Pro Previewgemini-3.1-pro-preview$2.00$12.00$0.20No

* Rate through 31 December 2026. Every figure is Google’s listed Standard paid-tier rate as the page gives it; the three Pro-class rows show the rate for prompts of 200,000 tokens or fewer.1

Price doesn’t follow version number. Gemini 3.5 Flash, which Google calls “Our earlier Flash model”, is the dearest Flash on the page.1 Its input rate is twice Gemini 3.8 Flash’s and its output rate is 2.4 times higher.1 Gemini 3.5 Flash-Lite also costs more than the older Gemini 3.1 Flash-Lite on both lines.1

The ranked field is the text input rate, and some models charge more for audio. Audio input costs $0.30 on Gemini 2.5 Flash-Lite, $0.50 on Gemini 3.1 Flash-Lite and $1.00 on both Gemini 2.5 Flash and Gemini 3 Flash Preview.1 Gemini 3.5 Flash-Lite and the two Robotics models list one input rate across text, image, video and audio.1 PDFs and other DOCUMENT tokens bill at the image token rate.1

Google’s output rates include thinking tokens.1 A model that reasons at length bills that reasoning at its output rate, which runs at 4 times the input rate on Gemini 2.5 Flash-Lite and about 8.3 times on Gemini 2.5 Flash.1 Every model in this table with a caching price charges one-tenth of its input rate for cached tokens. Storage then costs $1.00 per million tokens per hour on most of them and $4.50 on the two Pro models.1 The three current Flash models and the Robotics ER 2 Preview charge $0.50 per million tokens per hour for storage until the end of 2026.1

Which Gemini models charge more for prompts over 200,000 tokens?

Glamdring Research found three Gemini models with prompt-length price tiers on Google’s pricing page, checked 28 September 2026: Gemini 3.1 Pro Preview, Gemini 2.5 Pro and Gemini 2.5 Computer Use Preview.1 Each doubles its input rate and raises its output rate by half once a prompt exceeds 200,000 tokens.1 Every other model on the page lists one rate at any prompt length.1

ModelInput, prompts ≤200kInput, prompts >200kOutput, prompts ≤200kOutput, prompts >200kCached input, ≤200k / >200k
Gemini 3.1 Pro Preview$2.00$4.00$12.00$18.00$0.20 / $0.40
Gemini 2.5 Pro$1.25$2.50$10.00$15.00$0.125 / $0.25
Gemini 2.5 Computer Use Preview$1.25$2.50$10.00$15.00Not listed

All prices are USD per million tokens at Standard paid-tier rates.1

The tier is set by the prompt. Google’s Developer API tables label both the input and the output line by prompt size, as “prompts <= 200k tokens” and “prompts > 200k”.1 Google’s Agent Platform page states the rule in full: “If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates.”2 The Developer API page doesn’t repeat that sentence, so confirm the behaviour on your first long-context invoice.

Read that way, the threshold works as a cliff. One request with a 5,000-token answer costs:

Prompt sizeGemini 3.1 Pro PreviewGemini 2.5 Pro
200,000 tokens$0.46$0.30
200,001 tokens$0.89$0.58
500,000 tokens$2.09$1.33

Figures are Glamdring’s arithmetic on the listed Standard rates, rounded to the cent. They assume every token bills at the long-context rate once the prompt passes 200,000 tokens.1

The decision changes at the line. On Gemini 3.1 Pro Preview a 210,000-token prompt costs $0.93 with the same answer, and a 199,000-token prompt costs $0.46.1 Trimming 11,000 tokens of retrieved context halves that request’s cost. A retrieval system that packs context to the model’s limit should cap its prompts at 200,000 tokens unless the extra material changes the answer.

Batch and Flex keep the same two tiers at half price. Gemini 3.1 Pro Preview costs $1.00 and $2.00 for input and $6.00 and $9.00 for output under Batch.1

What do Gemini’s audio, image, video and embedding models cost?

Gemini’s audio, speech, image, video and embedding models are priced per million tokens like the language models. Google also gives a per-minute, per-image or per-second equivalent for most of them, and that is the unit to budget in.1

ModelInput (USD per 1M tokens)Output (USD per 1M tokens)Google’s stated equivalentFree tier
Gemini 3.8 Live, 3.8 Live Extended Thinking and 3.1 Flash Live Preview$0.75 text; $3.00 audio; $1.00 image/video$4.50 text; $12.00 audio$0.005/min audio in; $0.002/min image/video in; $0.018/min audio outYes
Gemini 3.5 Live Translate$3.50 audio$21.00 audioAbout $0.0368 per minuteYes
Gemini 3.5 Transcribe Live$3.50 audio$21.00 textAbout $0.009 per minute, blendedYes
Gemini 3.5 Transcribe$2.00 audio$12.00 textAbout $0.005 per minute, blendedYes
Gemini 2.5 Flash Native Audio (Live API)$0.50 text; $3.00 audio/video$2.00 text; $12.00 audioNot givenYes
Gemini 3.8 Flash TTS$0.50 text*$9.00 audio*$0.00225 per 10s of audio*Yes
Gemini 3.8 Flash-Lite TTS$0.50 text*$6.00 audio*$0.0015 per 10s of audio*Yes
Gemini 3.1 Flash TTS Preview$1.00 text$20.00 audioNot givenYes
Gemini 2.5 Flash Preview TTS$0.50 text$10.00 audioNot givenYes
Gemini 2.5 Pro Preview TTS$1.00 text$20.00 audioNot givenNo
Gemini 3.1 Flash Image (Nano Banana 2)$0.50 text/image$3 text and thinking; $60.00 images$0.067 per 1K image; $0.151 per 4K imageNo
Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite)$0.25$1.50 text and thinking; $30.00 images$0.0336 per 1K imageNo
Gemini 3 Pro Image (Nano Banana Pro)$2.00 text/image$12.00 text and thinking; $120.00 images$0.134 per 1K or 2K image; $0.24 per 4K imageNo
Gemini 2.5 Flash Image (Nano Banana)$0.30 text/image$30 per 1M image tokens$0.039 per imageNo
Gemini Omni Flash$1.50$9.00 text; $17.50 videoAbout $0.10 per second of 720p videoNo
Gemini Omni Flash Preview$1.50$9.00 text; $17.50 videoAbout $0.10 per second of 720p videoNo
Gemini Embedding 2$0.20 text; $0.45 image; $6.50 audio; $12.00 videoNot applicable$0.00012 per image; $0.00016 per second of audio; $0.00079 per frameYes

* Rate through 31 December 2026. All figures are Standard paid-tier rates as the page gives them.1

Three sections price outside tokens. Veo 3.1 Standard costs $0.40 a second at 720p and 1080p and $0.60 at 4k. Veo 3.1 Fast costs $0.10, $0.12 and $0.30 at those resolutions. Veo 3.1 Lite costs $0.05 at 720p and $0.08 at 1080p, with no 4k output.1 Lyria 3.5 costs $0.08 per full song. Lyria 3 Clip Preview costs $0.04 per 30-second clip and Lyria 3 Pro Preview $0.08 per song.1 None of the Veo or Lyria models has a free tier.1 Gemma 4 runs the other way: it is free of charge on the free tier and has no paid tier.1

The per-second units make video comparable. Gemini Omni Flash bills 5,792 output tokens per second of 720p video, so a 10-second clip costs about $1.01 at $17.50 per million tokens.1 The same 10 seconds cost $1.00 on Veo 3.1 Fast, $0.50 on Veo 3.1 Lite and $4.00 on Veo 3.1 Standard.1 Google describes Omni Flash as a video generation and editing model, so the price gap buys a different capability.1

What changes on 1 January 2027?

Seven Gemini models double their token prices on 1 January 2027 on Google’s pricing page, checked 28 September 2026. They are Gemini 3.8, 3.7 and 3.6 Flash, Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS, and the two Gemini Robotics ER 2 previews.1 Google’s Agent Platform page calls the current Flash rate “introductory pricing”.2

ModelInput nowInput from 1 Jan 2027Output nowOutput from 1 Jan 2027
Gemini 3.8, 3.7 and 3.6 Flash$0.75$1.50$3.75$7.50
Gemini Robotics ER 2 Preview and Streaming Preview$1.00$2.00$5.00$10.00
Gemini 3.8 Flash TTS$0.50$1.00$9.00$18.00
Gemini 3.8 Flash-Lite TTS$0.50$1.00$6.00$12.00

All prices are USD per million tokens at Standard paid-tier rates.1

Cached input and storage double with them. Gemini 3.8 Flash moves from $0.075 to $0.15 per million cached tokens and from $0.50 to $1.00 per million tokens per hour of storage.1

The change reorders the ranking. If no other price moves, the three current Flash models fall from a tie at sixth to a four-way tie with Gemini 3.5 Flash at $1.50 input.1 The Robotics previews reach $2.00, level with Gemini 3.1 Pro Preview.1 Gemini 3.8 Flash keeps a lower output rate than Gemini 3.5 Flash, at $7.50 against $9.00.1 Any 2027 budget built on September rates for these seven models is half the true figure.

How do Batch, Flex and Priority change the price?

Batch and Flex cost half of Standard on every Gemini language model that lists them, and Priority costs 1.8 times Standard, on Google’s pricing page checked 28 September 2026.1 Google lists “Batch API (50% cost reduction)” among paid-tier features.1

ModelStandard input / outputBatch and Flex input / outputPriority input / output
Gemini 2.5 Flash-Lite$0.10 / $0.40$0.05 / $0.20$0.18 / $0.72
Gemini 2.5 Flash$0.30 / $2.50$0.15 / $1.25$0.54 / $4.50
Gemini 3.8 Flash*$0.75 / $3.75$0.375 / $1.875$1.35 / $6.75
Gemini 2.5 Pro (≤200k)$1.25 / $10.00$0.625 / $5.00$2.25 / $18.00
Gemini 3.1 Pro Preview (≤200k)$2.00 / $12.00$1.00 / $6.00$3.60 / $21.60

* Rate through 31 December 2026. All prices are USD per million tokens as the page lists them.1

Cached-token rates follow no single rule under Batch and Flex. Gemini 3.8, 3.7 and 3.6 Flash and Gemini 3.1 Flash-Lite halve them.1 Gemini 3.1 Pro Preview, Gemini 3 Flash Preview and the 2.5 models keep the Standard cache rate.1 Google Search grounding costs the same on every tier. Google Maps grounding is unavailable under Batch and Flex on the 2.5 models.1

Each tier has its own capacity. Priority traffic defaults to 0.3 times the standard rate limit for each model and tier.3 Batch requests run under separate limits: 100 concurrent batch requests, a 2GB input file limit and a per-model cap on enqueued tokens.3 The captured pricing page gives Flex’s rates without its service terms. Confirm Flex latency and availability in Google’s documentation before routing production traffic to it.

What does the free tier include, and what does paid add?

The Gemini API free tier gives free input and output tokens on a limited set of models, and Google may use that content to improve its products.1 The paid tier adds higher rate limits, context caching, the Batch API and Google’s most advanced models, and paid content is excluded from product improvement.1

The free list covers Gemini 3.8, 3.7, 3.6 and 3.5 Flash, 3.5 and 3.1 Flash-Lite, 3 Flash Preview, 2.5 Pro, 2.5 Flash and 2.5 Flash-Lite.1 It also covers the Live, Transcribe, most TTS and Robotics models, Gemini Embedding 2 and Gemma 4.1 Gemini 3.1 Pro Preview, Gemini 2.5 Computer Use Preview, Omni Flash, all four image models, Gemini 2.5 Pro Preview TTS, Veo and Lyria are paid-only.1 Google AI Studio usage is free of charge in all available regions.1 It stays free until you link a paid API key.4

Moving to paid means linking a billing account and prepaying at least $5 in credits.4 The maximum prepaid balance is $5,000, and unused credits expire after 12 months.4 When the balance reaches $0, every API key on that billing account stops working and requests fail with an HTTP 402 error.4

Paid accounts climb three usage tiers:

Usage tierQualificationBilling tier capSpend rate limit per 10 minutes
FreeActive project or free trialN/AN/A
Tier 1Linked active billing account$250$10
Tier 2$100 paid and 3 days from first successful payment$2,000$50
Tier 3$1,000 paid and 30 days from first successful payment$20,000 - $100,000+$200

Qualifications and caps as Google’s rate-limits page lists them.3

The documentation doesn’t print free-tier request limits. Google shows active limits in AI Studio, and states that “Specified rate limits are not guaranteed and actual capacity may vary.”3 Limits are tighter on experimental and preview models.3

What does a real workload cost on the Gemini API?

In Glamdring Research’s worked example, a month of 300,000 text requests costs $150.00 on Gemini 2.5 Flash-Lite, $1,237.50 on Gemini 3.8 Flash and $3,600.00 on Gemini 3.1 Pro Preview. All three use Standard rates from Google’s pricing page, checked 28 September 2026.1 The workload is Glamdring’s illustration; no requests were run.

The estimate assumes:

  • 10,000 requests a day for 30 days.
  • 3,000 input tokens and 500 output tokens per request, thinking tokens included.
  • Text only, with no grounding and no caching, and every prompt under 200,000 tokens.

The month totals 900 million input tokens and 150 million output tokens.

ModelInput costOutput costStandard totalBatch total
Gemini 2.5 Flash-Lite$90.00$60.00$150.00$75.00
Gemini 3.1 Flash-Lite$225.00$225.00$450.00$225.00
Gemini 3.5 Flash-Lite$270.00$375.00$645.00$322.50
Gemini 2.5 Flash$270.00$375.00$645.00$322.50
Gemini 3 Flash Preview$450.00$450.00$900.00$450.00
Gemini 3.8 Flash, to 31 Dec 2026$675.00$562.50$1,237.50$618.75
Gemini 3.8 Flash, from 1 Jan 2027$1,350.00$1,125.00$2,475.00$1,237.50
Gemini 2.5 Pro$1,125.00$1,500.00$2,625.00$1,312.50
Gemini 3.5 Flash$1,350.00$1,350.00$2,700.00$1,350.00
Gemini 3.1 Pro Preview$1,800.00$1,800.00$3,600.00$1,800.00

Figures are Glamdring’s arithmetic on the listed rates.1

Output is one token in seven here, yet it makes up between 40 and 58 per cent of each bill. The output share is 45 per cent on Gemini 3.8 Flash.1 Cutting answer length or thinking effort moves the bill further than the same cut to the prompt.

Caching changes the input side. Suppose 2,000 of each prompt’s 3,000 tokens are a shared instruction block held in one cache all month. Gemini 3.8 Flash then costs $833.22, 32.7 per cent less.1 That total is $45.00 for 600 million cached tokens and $225.00 for 300 million uncached tokens. Output adds $562.50 for output and $0.72 to store 2,000 tokens for 720 hours at $0.50 per million per hour.1 The captured pricing page gives no minimum cache size or expiry rule, so confirm both before counting on the saving.

Grounding changes the other side. If one request in ten used Google Search grounding on Gemini 3.8 Flash, the first 5,000 searches a month would be free. The other 25,000 would cost at least $350.00 at $14 per 1,000.1 The floor is “at least” because Google charges for each search query, and one request “may result in one or more queries to Google Search”.1 Gemini 2.5 models price grounding differently: 1,500 free requests a day, then $35 per 1,000 grounded prompts.1

Why do other published Gemini prices disagree?

Other Gemini price lists differ from Google’s Developer API page for four reasons. They quote Google Cloud’s Agent Platform, the 2027 rate, a discounted tier or a model the pricing page no longer lists.

Google Cloud’s Agent Platform is a separate price list. The Developer API page says its prices “may differ” from Agent Platform prices.1 Agent Platform lists Gemini 3.8 Flash at the same $0.75 and $3.75 in its Global region, and at $0.825 and $4.125 in non-global regions, 1.1 times higher.2 Some Agent Platform tables already show Gemini 3.8 Flash at $1.50 and $7.50, the standard rate Google says applies from 1 January 2027.2

Discounted rates get quoted as headlines. Some price lists give $0.05 per million input tokens as the cheapest Gemini rate. The same figure matches Gemini 2.5 Flash-Lite’s Batch and Flex input rate, and its Standard rate is $0.10.1

Older model names linger. Google’s rate-limits page still lists Gemini 2.0 Flash, 2.0 Flash Lite and 2.0 Flash Image in its Batch tables.3 The pricing page carries no price for any of them.1

Google’s own page carries two embedding prices. File search is charged for embeddings at $0.15 per million tokens. Gemini Embedding 2 lists text input at $0.20.1 The page doesn’t explain the gap, so price a file-search workload from the file-search line.

What the ranking cannot tell you

The ranking lists price per token. A finished task’s cost also depends on how many tokens the model spends, and output rates include thinking tokens.1 Two models at the same rate can bill very differently for the same job. Run one real workload on two models and compare the usage lines before choosing on list price.

The ranking covers only the Gemini Developer API page as captured on 28 September 2026.1 It doesn’t show Enterprise-tier volume discounts or provisioned throughput, which Google offers through its Enterprise plan.1 It doesn’t show tax, currency conversion or the free-tier rate limits, which live in AI Studio.3 Tool use adds its own charges. Code execution bills as tokens, URL context bills as input tokens and grounding bills per search.1

Confirm the rates on Google’s live page before committing a budget. More of our work on model serving and compute sits in the model infrastructure category. The ranking changes when Google next revises its price list, and on 1 January 2027 at the latest; the revision will carry a change note. Subscribe to our research to receive it.

Frequently asked questions

Is the Gemini API still free?

Yes, for a limited set of models. The free tier includes free input and output tokens on models such as Gemini 3.8 Flash, Gemini 2.5 Flash-Lite and Gemini 2.5 Pro, within free-tier rate limits, checked 28 September 2026.1 Gemini 3.1 Pro Preview, the image models, Omni Flash, Veo and Lyria have no free tier.1 Google may use free-tier content to improve its products.1

How do you use the Gemini API for free?

Use the project and API key that Google AI Studio creates for new users, and leave billing unlinked. New accounts begin on the Free Tier, with access to certain models up to their free-tier rate limits.4 A paid project returns to the free tier when you disable billing on it.4 AI Studio shows the active limits for each project.3

Does the Google Cloud free trial cover Gemini API usage?

No. From March 2026, Google excludes Gemini API usage costs from the $300 Google Cloud Free Trial.4 Welcome and free-trial credits can’t pay for the Gemini API or AI Studio.4 Free use comes from the Gemini API free tier and from AI Studio, which is free of charge in available regions.1

Sources checked

  1. Gemini Developer API pricingChecked September 28, 2026
  2. Agent Platform PricingChecked September 28, 2026
  3. Rate limits | Gemini APIChecked September 28, 2026

Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.