Groq API pricing: Groq model list, September 2026
For production text work at a public price, Groq offers two models: GPT OSS 20B at $0.075 input and $0.30 output per 1M tokens, and GPT OSS 120B at $0.15 and $0.60.

Groq’s cheapest token-priced model is Llama Prompt Guard 2 22M, at $0.03 per million input tokens and $0.03 per million output tokens, on Groq’s Supported Models list checked 28 September 2026.1 The dearest is Qwen3.8 27B, at $0.80 input and $4.00 output.1 For production text work, only GPT OSS 20B ($0.075 input, $0.30 output) and GPT OSS 120B ($0.15 input, $0.60 output) carry a public price.1
TL;DR
- Groq lists 13 models. Six are priced per million tokens, two per hour of audio, two per million characters, and three show “Contact Sales” instead of a price.1
- GPT OSS 120B costs exactly twice GPT OSS 20B, on input and on output.1
- Qwen3.8 27B is the only token-priced model above $1 per million output tokens, and it’s a preview model Groq says may be discontinued at short notice.1
- Llama 3.3 70B shows no self-serve price on the checked list, while a third-party comparison still quotes $0.59 input and $0.79 output for it on Groq.1,2
- Requesty lists Groq’s GPT OSS 20B at $0.10 input and $0.50 output, 33% and 67% above Groq’s own list, before its 5% pay-as-you-go fee.1,3
- Groq’s free tier is capped by rate limits, not a bill. The captured evidence gives exact free limits for the Whisper models only.4
How this ranking was built
Every model’s input and output price from Groq’s published model list, checked 28 September 2026, is in the two tables below. The evidence rules behind that check are set out in our research methodology. The source is Groq’s Supported Models page in its GroqCloud documentation, captured at 15:06 AEST on 28 September 2026.1
The ranked field is Groq’s Price per 1M tokens column. Models are ordered by their output rate, cheapest first, and the input rate breaks ties.1
Row counts at each step:
- Listed: 13 models, 6 in Groq’s production section and 7 in its preview section. The deprecated section names no models and points to a separate deprecations page.1
- Ranked: the 6 models with a per-token price.
- Listed separately: the 7 models priced per hour of audio, per million characters, or by quote. Those units can’t sit on the same scale as tokens.
- Merged or excluded: none.
Prices appear as Groq prints them, with a dollar sign. The page doesn’t name the currency or say whether tax is included.1 These are Groq’s own list prices: a company statement of what it charges, not a measured market price. More pricing and serving research sits in the model infrastructure category.
What does each Groq model cost per million tokens?
GPT OSS 20B is the cheapest model on Groq’s list with a 131,072-token context window, at $0.075 per million input tokens and $0.30 per million output tokens.1 Only the two Prompt Guard models cost less, and each accepts just 512 tokens of context.1
| Rank | Model | Model ID | Status | Input ($ per 1M tokens) | Output ($ per 1M tokens) | Context window (tokens) | Developer-plan limits |
|---|---|---|---|---|---|---|---|
| 1 | Llama Prompt Guard 2 22M | meta-llama/llama-prompt-guard-2-22m | Preview | $0.03 | $0.03 | 512 | 30K TPM, 100 RPM |
| 2 | Prompt Guard 2 86M | meta-llama/llama-prompt-guard-2-86m | Preview | $0.04 | $0.04 | 512 | 30K TPM, 100 RPM |
| 3= | GPT OSS 20B | openai/gpt-oss-20b | Production | $0.075 | $0.30 | 131,072 | 250K TPM, 1K RPM |
| 3= | Safety GPT OSS 20B | openai/gpt-oss-safeguard-20b | Preview | $0.075 | $0.30 | 131,072 | 150K TPM, 1K RPM |
| 5 | GPT OSS 120B | openai/gpt-oss-120b | Production | $0.15 | $0.60 | 131,072 | 250K TPM, 1K RPM |
| 6 | Qwen3.8 27B | qwen/qwen3.8-27b | Preview | $0.80 | $4.00 | 131,072 | 250K TPM, 1K RPM |
Source: Groq, Supported Models, checked 28 September 2026. TPM is tokens per minute and RPM is requests per minute, as Groq abbreviates them.1
Three readings change a budget.
The GPT OSS pair scales cleanly. GPT OSS 120B costs twice GPT OSS 20B on both rates, and on both models output costs four times input.1 A workload that tolerates the smaller model halves its token bill.
Qwen3.8 27B sits in a different price band. Its output rate is 6.67 times GPT OSS 120B’s, its input rate 5.33 times, and its output costs five times its input.1 It also caps a single completion at 16,384 tokens, against 65,536 for both GPT OSS models.1
Safety GPT OSS 20B costs the same as GPT OSS 20B but carries a lower developer-plan limit: 150K tokens per minute against 250K.1
A worked example makes the spread concrete. One million input tokens plus 100,000 output tokens costs $0.105 on GPT OSS 20B, $0.21 on GPT OSS 120B and $1.20 on Qwen3.8 27B at list rates.1 The estimate assumes no caching, batch or flex discount, because none appears on the model list. Token rates are one input to the bill; our explainer on what controls AI inference cost covers the drivers that sit behind a per-token price.
Which Groq models are priced per hour, per character or by quote?
Seven of Groq’s 13 models have no per-token price. The two Whisper models are billed per hour of audio, the two Orpheus models per million characters, and three models show “Contact Sales”.1
| Model | Model ID | Status | Listed price | Unit |
|---|---|---|---|---|
| Whisper Large V3 Turbo | whisper-large-v3-turbo | Production | $0.04 | per hour |
| Whisper | whisper-large-v3 | Production | $0.111 | per hour |
| Canopy Labs Orpheus V1 English | canopylabs/orpheus-v1-english | Preview | $22.00 | per 1M characters |
| Canopy Labs Orpheus Arabic Saudi | canopylabs/orpheus-arabic-saudi | Preview | $40.00 | per 1M characters |
| Llama 3.1 8B | llama-3.1-8b-instant | Production, Enterprise | Contact Sales | none listed |
| Llama 3.3 70B | llama-3.3-70b-versatile | Production, Enterprise | Contact Sales | none listed |
| MiniMax M2.7 | minimaxai/minimax-m2.7 | Preview, Enterprise | Contact Sales | none listed |
Source: Groq, Supported Models, checked 28 September 2026.1
Whisper large-v3 costs 2.78 times the turbo model per hour.1 By Groq’s own figures, as reported by VoiceBoard, the dearer model buys a lower word error rate, 10.3% against 12%, and translation, which turbo doesn’t offer.4 VoiceBoard also reports that a request shorter than 10 seconds is billed as 10 seconds, so short clips cost more per second than the hourly rate implies.4
The Arabic Saudi Orpheus voice costs 1.82 times the English one per million characters.1
The quote-only rows matter most. Both Llama models sit in Groq’s production section, yet neither shows a price or a rate limit.1 A team that wants Llama on Groq starts with a sales conversation, not a price sheet.
Is there a free tier on Groq’s API?
Yes. Groq’s free tier applies from the first request and is capped by rate limits rather than by a bill: past the limit, requests stop until the window resets.4 Moving to the paid tier is a deliberate step.4
The exact free limits in the captured evidence cover the Whisper models only. VoiceBoard read them from Groq’s rate-limits page on 10 September 2026:4
- Requests: 20 a minute, 2,000 a day.4
- Audio: 7,200 seconds an hour, 28,800 seconds a day.4
- File size: 25 MB on the free tier, 100 MB on the paid Dev tier.4
The 100 MB figure matches the maximum file size on Groq’s own model list, where the rate-limit column is labelled “Developer plan”.1 That column shows paid-plan limits, not free ones.
Free-tier limits for the text models weren’t in the evidence captured for this page. Groq publishes its free limits on its rate-limits page, which its model list links to.1,4 Confirm the current free limit there before sizing a prototype on it.
How do Groq’s prices compare with the same model elsewhere?
Where the captured evidence covers the same model elsewhere, Groq’s own list is cheaper. GPT OSS 20B costs $0.075 input and $0.30 output on Groq’s list, against $0.10 and $0.50 in Requesty’s catalogue listing for Groq’s deployment of the same model.1,3
Requesty sells access to many providers’ models through one key, and its id for this listing calls Groq’s own deployment directly.3 It describes $0.10 and $0.50 as “the upstream provider rates”, updated 27 September 2026, and adds 5% on pay-as-you-go.3 Those rates sit 33% above Groq’s list on input and 67% above on output. Requesty prices one million input tokens plus 100,000 output tokens at $0.15.3 Groq’s list rates give $0.105 for the same workload, so Requesty’s figure is 43% higher, or 50% higher once the 5% fee is added.1 Neither page explains the gap.
For GPT OSS 120B, Requesty’s listing shows $0.15 per million tokens, which matches Groq’s input rate.3 Its output rate for that model wasn’t captured.
Speech-to-text shows the widest spread. OpenAI’s whisper-1 endpoint costs $0.006 a minute, or $0.36 an hour, as VoiceBoard read it from OpenAI’s pricing page on 10 September 2026.4 That’s 3.24 times Groq’s whisper-large-v3 rate and 9 times its turbo rate.1 The captured evidence doesn’t show that whisper-1 runs the same weights as whisper-large-v3, so treat this as the same model family, not an identical deployment.
For Qwen3.8 27B, Safety GPT OSS 20B, both Prompt Guard models and both Orpheus voices, no price from another host was captured. This page makes no claim about them.
One boundary applies to any cross-provider comparison. Groq serves open-weight models on its own chips, while OpenAI and Anthropic sell their own proprietary models, so a price gap against those models is a gap between different products.2
Why do other published Groq prices disagree with the list?
Third-party pages quote Groq prices that don’t match the list checked on 28 September 2026, and none of the captured pages says why.1
Flowagenz’s cost-per-conversation comparison prices Llama 3.3 70B on Groq at $0.59 input and $0.79 output per million tokens.2 On the checked list, that model shows “Contact Sales” and an Enterprise tag.1 The same comparison gives Groq’s typical range as $0.05 to $1.00 input and $0.08 to $3.00 output.2 Qwen3.8 27B’s $4.00 output rate sits outside that range.1
Requesty’s GPT OSS 20B rates, covered above, are the second mismatch.3
The decision changes when a budget rests on a quoted figure. Price against Groq’s own page on the day, and confirm it again in the current quote before committing spend.
Which priced Groq models are cleared for production?
Four models have both a public price and production status: GPT OSS 20B, GPT OSS 120B and the two Whisper models.1 Every other priced model is in Groq’s preview section.
Groq says preview models are intended for evaluation only, should not be used in production, and may be discontinued at short notice.1 That covers Qwen3.8 27B, Safety GPT OSS 20B, both Prompt Guard models and both Orpheus voices.1
The constraint is status, not price. A production text workload that needs a public rate on Groq has two choices, GPT OSS 20B and GPT OSS 120B. Everything else is either preview or quote-only.
What the ranking cannot tell you
- Other service tiers and caching. Groq’s documentation has pages for Prompt Caching, Flex Processing, Batch Processing and a Performance Tier.1 The model list prints one price per model, and those tiers’ rates weren’t captured, so the tables don’t show them.
- Enterprise prices. Three models are quote-only. Their terms aren’t public.
- Currency and tax. The page prints a dollar sign without naming the currency or tax treatment.1
- Speed and quality. The speed column, such as 1,000 tokens a second for GPT OSS 20B and 500 for GPT OSS 120B, is Groq’s own figure.1 We haven’t measured it, and the list says nothing about output quality.
- Cost per task. Cost per conversation depends on turn count, context growth and caching at least as much as on the per-token rate, as Flowagenz’s analysis sets out.2 Use one real workflow to test the claim.
- What changed. This is the first capture of the list for this page, so it can’t show which prices moved or when.
The ranking refreshes when Groq changes its Supported Models page. New price checks go out in our email briefing.
Frequently asked questions
Is the Groq API key free to use?
Yes. A Groq key starts working immediately, and the free tier applies from the first request.4 Use is capped by rate limits: past a limit, requests stop until the window resets rather than running up a bill.4 Paid list prices apply once you move to a paid plan on purpose.4
How much does GPT cost on Groq’s API?
Groq serves OpenAI’s open-weight GPT OSS models, not OpenAI’s proprietary GPT models.2 On the list checked 28 September 2026, GPT OSS 20B costs $0.075 input and $0.30 output per million tokens, and GPT OSS 120B costs $0.15 and $0.60.1 For OpenAI’s own frontier models, Flowagenz puts typical rates at $2.50 to $5.00 input and $15 to $30 output per million tokens.2 Confirm OpenAI’s figure on OpenAI’s pricing page.
Does a router add a fee on top of Groq’s price?
Requesty adds 5% to its listed rates on pay-as-you-go, or 0% if you bring your own provider keys, with no per-request fee.3 Its listed rate for GPT OSS 20B on Groq is already above Groq’s own list, so compare the router’s base rate as well as its fee.1,3
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


