Groq vs Cerebras: price pages and speed claims, Sep 2026
Groq lists GPT-OSS 120B at $0.15 input and $0.60 output per million tokens, against Cerebras's $0.35 and $0.75. Cerebras charges less on Qwen 3.8 27B below about 13.2 input tokens per output token. Its production capacity is quote-only.

Groq costs less per million tokens on GPT-OSS 120B: $0.15 for input and $0.60 for output, against $0.35 and $0.75 on Cerebras.1,2 On Qwen 3.8 27B, the only other model both companies price publicly, Cerebras costs less below about 13.2 input tokens per output token. Its output rate is $1.49 against Groq’s $4.00.1,2 Cerebras claims more speed on both: about 3,000 tokens per second on GPT-OSS 120B against Groq’s 500.1,2 Those speed figures are vendor claims, checked 28 September 2026.
TL;DR
- Groq publishes a production price for GPT-OSS 120B. Cerebras’s production capacity is quote-only. Its published exploration-tier rate is 71% higher at the 3:1 input-to-output mix used in its own comparison.1,2,3
- Qwen 3.8 27B reverses the price answer. Cerebras costs less unless a workload sends more than about 13.2 input tokens per output token. Groq lists the model as a preview for evaluation only.1
- Cerebras claims 6 times Groq’s speed on GPT-OSS 120B and about 4.1 times on Qwen 3.8 27B. Both figures come from the companies’ own pages. Cerebras says its performance comparisons rest on “third-party benchmarking or internal testing”.2
- Cerebras’s public prices sit in a Developer tier for exploration workloads, with a free $5 credit. Production capacity is quote-only.2 Groq’s page shows Developer plan rate limits but states no free-tier terms.1
How these prices were checked
Glamdring Research compared each company’s own page, captured on 28 September 2026. Groq’s figures come from its GroqCloud Supported Models page, using the “Price per 1M tokens” and “Speed (T/sec)” columns.1 Cerebras’s figures come from the Developer Tier Pricing table on its pricing page, using the “Input”, “Output” and “Speed” columns.2 Both pages show prices with a dollar sign. Neither names a currency code or a tax treatment.
The inclusion rule was a text model with a public per-token price on both pages. Groq’s page lists 13 models across its production and preview tables. Six of them are text models with a public per-token price. Cerebras’s table prices two. Two models appear on both: GPT-OSS 120B and Qwen 3.8 27B. Groq lists Llama 3.1 8B and Llama 3.3 70B as Enterprise models priced through “Contact Sales”, so they drop out.1
The blended rates below are our own arithmetic. They weight input three to one against output because that’s the ratio Cerebras used when it compared its prices with Groq’s.3 The research methodology sets out how Glamdring records and recomputes vendor figures, and the comparison sits within our model infrastructure coverage.
Which is cheaper per million tokens on the models both host?
Groq is cheaper on GPT-OSS 120B at every token mix, and Cerebras is cheaper on Qwen 3.8 27B at an assumed 3:1 mix. On GPT-OSS 120B, Groq charges $0.15 per million input tokens and $0.60 per million output tokens; Cerebras charges $0.35 and $0.75.1,2 On Qwen 3.8 27B, Groq charges $0.80 and $4.00 while Cerebras charges $0.99 and $1.49.1,2 Both sets of prices were checked on 28 September 2026.
| Model | Groq input ($/1M) | Groq output ($/1M) | Cerebras input ($/1M) | Cerebras output ($/1M) | Groq blended 3:1 ($/1M) | Cerebras blended 3:1 ($/1M) | Groq model status |
|---|---|---|---|---|---|---|---|
| GPT-OSS 120B | $0.15 | $0.60 | $0.35 | $0.75 | $0.2625 | $0.45 | Production |
| Qwen 3.8 27B | $0.80 | $4.00 | $0.99 | $1.49 | $1.60 | $1.115 | Preview |
Prices as each page shows them, checked 28 September 2026.1,2 Blended rates are computed: three parts input price plus one part output price, divided by four.
GPT-OSS 120B gives Cerebras no opening. Its input rate is 2.33 times Groq’s and its output rate is 1.25 times Groq’s, so no mix of tokens brings the two bills level. A job of 3 million input tokens and 1 million output tokens costs $1.05 on Groq and $1.80 on Cerebras.
Cerebras’s own September 2025 comparison said it was “priced similarly to Groq or up to ~50% higher, assuming a 3:1 input : output token ratio”.3 At the rates both companies published on 28 September 2026, the GPT-OSS 120B gap at that ratio is 71%. The older statement predates the current price pages, so the current pages govern.
Groq also lists a smaller sibling, GPT-OSS 20B, at $0.075 input and $0.30 output with a claimed 1,000 tokens per second.1 Cerebras’s price table has no equivalent. A buyer who can accept the smaller model gets a lower rate on Groq with no Cerebras price to compare against.
When does Cerebras cost less than Groq?
Cerebras costs less than Groq on Qwen 3.8 27B whenever a workload sends fewer than about 13.2 input tokens for each output token. Groq charges $0.80 per million input tokens and $4.00 per million output tokens. Cerebras charges $0.99 and $1.49.1,2 Cerebras’s input premium is $0.19 per million. Groq’s output premium is $2.51 per million. The bills meet where 0.19 times the input volume equals 2.51 times the output volume, which is 13.2 to 1.
At 3:1, Cerebras’s blended rate is $1.115 per million tokens against Groq’s $1.60, so Groq costs 43% more. At 20 input tokens per output token, Groq’s blended rate is about $0.95 against Cerebras’s $1.01. The cheaper endpoint depends on the workload’s input-to-output ratio.
The status of each price matters as much as its size. Groq files Qwen 3.8 27B under preview models, which it says “are intended for evaluation purposes only and should not be used in production environments as they may be discontinued at short notice”.1 Cerebras labels its Developer tier “Exploration workloads” and reserves “Production-ready capacity” for its Enterprise tier, which is priced by quote.2 Neither public Qwen 3.8 27B price is a production price.
The Cerebras page also disagrees with itself. Its tier table lists the Developer models as “gpt-oss 120b, gemma-4-31b”, while its Developer price table lists GPT OSS 120B and Qwen 3.8 27B and gives no Gemma price.2 Confirm Qwen 3.8 27B availability in the Cerebras console before planning on that rate. For the wider set of levers that move an inference bill, see our analysis of what controls AI inference cost.
Which claims more speed, Groq or Cerebras?
Cerebras claims more speed on both models the two companies price: about 3,000 tokens per second on GPT-OSS 120B against Groq’s 500, and about 1,850 on Qwen 3.8 27B against Groq’s 450.1,2 These are vendor claims from each company’s own page, checked 28 September 2026. Taken at face value, Cerebras claims 6 times Groq’s output speed on GPT-OSS 120B and about 4.1 times on Qwen 3.8 27B. Neither page publishes the prompt length, concurrency or region behind its figure.
| Model | Groq claimed speed (tokens/s) | Cerebras claimed speed (tokens/s) | Cerebras claim as multiple of Groq claim |
|---|---|---|---|
| GPT-OSS 120B | 500 | ~3,000 | 6.0 |
| Qwen 3.8 27B | 450 | ~1,850 | 4.1 |
Vendor claims as published, checked 28 September 2026.1,2 Multiples are computed.
At the claimed rates, a 1,000-token answer from GPT-OSS 120B streams in about 2 seconds on Groq and about a third of a second on Cerebras.
Cerebras qualifies its own numbers. Its pricing page says “Performance comparisons are based on third-party benchmarking or internal testing” and that observed speed “may vary depending on workload, configuration, date and models being tested”.2 Groq’s models page gives no method for its speed column.1
Third-party figures point the same way but spread widely. Cerebras’s September 2025 comparison cited Artificial Analysis at about 3,000 tokens per second for Cerebras and about 493 for Groq on GPT-OSS 120B.3 The figure reaches us through Cerebras, and the original benchmark wasn’t captured for this comparison. GMI Cloud, a GPU cloud company, put Groq at 476 tokens per second on the same model in May 2026.4 Spheron, which rents GPUs, gave ranges of about 478 to 493 for Groq and about 1,700 to 3,000 for Cerebras in August 2026.5 Every captured source puts Cerebras ahead on throughput. None of them measured both services under the same published conditions.
Throughput is also only one speed. GMI Cloud reports time to first token under 100 milliseconds on Groq and 80 to 150 milliseconds on Cerebras, without naming its test.4 Neither company’s captured page publishes a first-token figure. GMI Cloud treats voice assistants as the case where first-token time decides the choice.4
Does Cerebras’s speed claim change the cost answer?
Glamdring Research’s reading is that the speed claim changes the cost answer only when time has a price. A per-token bill doesn’t reward speed. A million output tokens on GPT-OSS 120B costs $0.60 on Groq and $0.75 on Cerebras however fast they arrive.1,2 Cerebras’s claim of “up to a 6x price-performance advantage over Groq” folds its speed claim into the price comparison.3
Spheron reads the claim the same way. It says the claim is “about throughput per dollar spent generating tokens fast enough to matter for a workload, not about the sticker price per token being lower”.5 A team “paying purely by the token with no time pressure on the workload” sees none of the throughput advantage in its bill.5 Where wall-clock time costs money, Spheron argues Cerebras “can come out ahead even at a higher price per token, because it finishes the job faster”.5 Its examples include an agent loop with a latency budget and a batch job that must finish in a set time. Work with no deadline sees only the per-token rate.
The decision changes when a team can put a number on a second of waiting. Without that number, the per-token rate decides, and on GPT-OSS 120B the per-token rate favours Groq. Use one real workflow to test the claim: run the same prompts through both APIs, then record the bill and the completion time side by side.
What do the plans and free tiers include?
Cerebras publishes two tiers: a self-serve Developer tier for exploration with a free $5 credit, and an Enterprise tier priced by quote.2 Groq’s models page shows per-model rate limits under a Developer plan and routes some models to Enterprise pricing through sales.1 The captured Groq page states no free-tier terms or starting credit.
| Plan detail | Groq (models page) | Cerebras (pricing page) |
|---|---|---|
| Self-serve tier | Developer plan; rate limits shown per model | Developer, “Exploration workloads”; self-serve pay-as-you-go |
| Starting credit | Not stated on the captured page | Free $5 credit to start |
| Published self-serve limits | GPT-OSS 120B and Qwen 3.8 27B: 250K tokens per minute, 1K requests per minute | Not published on the pricing page |
| Enterprise | Llama 3.1 8B, Llama 3.3 70B and MiniMax M2.7 priced via “Contact Sales” | “Contact us for a quote”; all preview and production models |
| Enterprise-only features | Not listed on the models page | Production-ready capacity; best priority, rate limits and latency; custom model weights; fine-tuning and training services |
| Developer support | Not stated on the captured page | Community Discord |
| Other service options | Documentation lists Performance Tier, Flex Processing and Batch Processing | Partner access through AWS Marketplace, OpenRouter, Hugging Face and Vercel |
Plan terms as each page showed them on 28 September 2026.1,2
The two structures put the production line in different places. Groq publishes self-serve prices for its production models, including GPT-OSS 120B.1 Cerebras publishes self-serve prices only for exploration, and every production price is a quote.2 On the 28 September pages, a team that wants a production price without a sales call gets one from Groq.
Cerebras’s tier structure has moved recently. In August 2026, Spheron described a Cerebras “Free Trial tier with $5 in credits” plus “a self-serve Developer tier starting at $10”.5 The 28 September page shows a single pay-as-you-go Developer tier with the $5 credit folded in.2 Groq’s Performance Tier, Flex Processing and Batch Processing pages appear in its documentation menu, but their terms weren’t captured here.1 Our Groq API pricing breakdown covers Groq’s own price list on its own terms.
What this comparison cannot tell you
The comparison can’t tell you whether the same model name returns the same quality on both services. Cerebras’s comparison says “most models running on Groq trade off accuracy for speed through quantization down to 8-bit precision” while Cerebras “supports 16-bit precision natively in hardware”.3 Cerebras is describing a rival there. Groq’s captured page doesn’t state precision. Test output quality on your own prompts before treating the two GPT-OSS 120B endpoints as interchangeable.
It can’t tell you speed under load. Both speed columns give one figure per model. GMI Cloud says the gap against GPUs “is largest at low concurrency and smallest at high concurrency”.4 Spheron argues that above roughly 8 simultaneous requests, a rented H100 “produces more tokens per dollar than either Groq or Cerebras”.5 Both companies sell GPU capacity, and that assumption travels with their conclusions.
It can’t tell you Enterprise prices. Groq’s Llama models and every Cerebras production workload sit behind a quote.1,2 Confirm any production rate in a current quote.
The verdict changes if either company reprices GPT-OSS 120B, Groq moves Qwen 3.8 27B into production, Cerebras publishes production prices, or an independent benchmark measures both services under the same published conditions. Glamdring’s email briefing carries the next check when these pages change.
Frequently asked questions
Is Groq cheaper than Cerebras on Llama 3.3 70B?
Neither company’s captured page shows a self-serve Llama 3.3 70B price on 28 September 2026. Groq lists it as an Enterprise model priced through “Contact Sales”, with a claimed 280 tokens per second.1 Cerebras’s Developer price table doesn’t list it.2 Spheron’s August 2026 comparison quoted $0.59 and $0.79 per million tokens for Groq against about $0.85 and $1.20 for Cerebras. It took the Cerebras rate from a third-party cost calculator.5 Those figures no longer match either company’s own page, so ask both for a quote.
Can you fine-tune models on Groq or Cerebras?
Cerebras offers fine-tuning only in its Enterprise tier, which lists “Customize model weights” and “Fine-tuning and training services”.2 Groq’s documentation includes a LoRA Inference page, but its terms weren’t captured here.1 Spheron reported in August 2026 that Groq “gates LoRA fine-tuning entirely behind its Enterprise tier and a sales request”.5 On either service, plan for a sales conversation before a fine-tuned model reaches production.
Is Cerebras faster than Nvidia GPUs?
Cerebras says it is, and attributes its performance comparisons to “third-party benchmarking or internal testing”.2 The same page says speed improvements against GPU-based systems vary with workload, configuration, date and model. GMI Cloud reports that the gap is widest at low concurrency and narrows as concurrency rises.4 Spheron puts the crossover at roughly 8 simultaneous requests, above which a rented H100 yields more tokens per dollar.5 Both of those companies sell GPU capacity.
Who competes with Cerebras?
Groq is the rival Cerebras chose for its own head-to-head comparison, published in September 2025.3 Spheron names SambaNova’s SN40L as “a third SRAM-and-dataflow-style challenger”.5
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


