Groq competitors: gpt-oss-120b prices, September 2026
SambaNova lists the lowest gpt-oss-120b output figure, 0.59 USD per million tokens. Total-bill comparisons depend on input mix, currency and billing mode; Groq and Together AI have lower listed input figures, and Cerebras's self-serve rate is an exploration offer.

SambaNova lists the lowest output figure among the four quoted gpt-oss-120b price rows: 0.59 USD per million output tokens, checked 28 September 2026.1 Groq lists $0.60,2 Together AI also lists $0.603 and Cerebras lists $0.75.4 Fireworks AI’s pricing landing page prints no per-token rate for the model, so it isn’t ranked in this dated page comparison.5 SambaNova’s one-cent output advantage must be weighed against its higher input price.
TL;DR
- Three of the four ranked providers sit within one cent of each other on the listed output price per million tokens. Cerebras is the outlier at $0.75, 25% above Groq’s $0.60.4,2
- On the listed numbers, SambaNova’s output advantage does not guarantee a lower total. Its input price is 0.22 USD per million tokens against $0.15 at Groq and Together AI.1,2,3 Assuming the same dollar currency and listed billing rates, SambaNova costs less only above seven output tokens per input token.
- Only Groq and Cerebras print speed figures beside their prices. Both are vendor claims, stated on each company’s own terms.2,4
- Cerebras ($5) and Fireworks AI ($1) state free starting credit on their pricing pages. The Groq, SambaNova and Together AI pages checked state none.4,5
- Llama 3.3 70B no longer works as a shared yardstick. Groq lists it as an Enterprise model with a contact-sales price.2
How this ranking was built
The ranking uses one field: the output price per 1M tokens that each provider prints for gpt-oss-120b, checked 28 September 2026. Lower ranks higher. Figures appear exactly as printed. SambaNova names the field Output (per 1M tokens) and labels its figures USD.1 The other pages print a dollar sign without naming a currency, and none of the five states whether tax is included. Groq publishes per-token prices in the model table of its models documentation, so that table is Groq’s source; our Groq API pricing breakdown covers its full rate card.2
This is a comparison of printed price numbers, not currency-converted invoices or matched production contracts. The worked bills below assume the listed dollar amounts use the same currency and billing mode. Confirm those terms before treating the arithmetic as an actual cost comparison.
Together’s September capture also contains a “Batch API price” label. A separate check of its official model page on 1 October confirms the same $0.15 input and $0.60 output figures beside Serverless and Dedicated deployment options, rather than a batch-specific rate.6 That later check clarifies the listing; it is not backdated as a September capture.
Row counts at each step:
- Five providers considered. Cerebras and SambaNova appear in published chip comparisons against Groq.7 Together AI and Fireworks AI appear in a published list of Groq alternatives.8
- Five pricing pages captured on 28 September 2026.
- Two candidate shared models tested. gpt-oss-120b carries a printed per-token price on four of the five pages. Llama 3.3 70B carries one on two.
- One exclusion. Fireworks AI’s pricing page sends readers to its documentation for serverless rates and prints none for gpt-oss-120b.5
- Four providers ranked. Ties share a rank: Groq and Together AI both list $0.60.
Discounted tiers sit outside the rule. Groq’s documentation lists Flex Processing and Batch Processing tiers; the ranking uses the standard model-table price.2 Proprietary-model APIs that some lists name as Groq alternatives, such as OpenAI and Anthropic, fall outside it too, because the ranking compares one open model.8 The general research standard is set out on the methodology page.
Which Groq competitor lists the lowest gpt-oss-120b price?
Glamdring Research’s check of four pricing pages ranks SambaNova first for gpt-oss-120b, at 0.59 USD per million output tokens, checked 28 September 2026.1 Groq and Together AI tie second at $0.60.2,3 Cerebras is fourth at $0.75.4 The spread from first to last is 16 cents per million output tokens.
| Rank | Provider | Model as listed | Output, $ per 1M tokens | Input, $ per 1M tokens | Listing status | Speed figure on page (vendor claim) |
|---|---|---|---|---|---|---|
| 1 | SambaNova | gpt-oss-120b | 0.59 USD | 0.22 USD | Pricing table | None printed |
| 2= | Groq | openai/gpt-oss-120b | $0.60 | $0.15 | Production model | 500 tokens/s |
| 2= | Together AI | gpt-oss-120B | $0.60 | $0.15 | Serverless inference | None printed |
| 4 | Cerebras | GPT OSS 120B | $0.75 | $0.35 | Developer tier | ~3000 tokens/s |
| Not ranked | Fireworks AI | Not printed | Not printed | Not printed | Serverless; rates in documentation | None printed |
Sources: SambaNova1, Groq2, Together AI3, Cerebras4, Fireworks AI5. All checked 28 September 2026.
SambaNova undercuts Groq by one cent, about 1.7%. Together AI matches Groq on both input and output. Cerebras costs more on both fields.
Listing status changes how to read two rows. Groq lists gpt-oss-120b as a production model, with Developer-plan rate limits of 250K tokens per minute and 1K requests per minute.2 Cerebras sells the model on its Developer tier, which it labels for exploration workloads. Production-ready capacity sits on its Enterprise tier, priced by quote.4 Cerebras’s $0.75 is a self-serve exploration price. A production buyer there would be comparing Groq’s listed rate against an unpublished Enterprise quote.
Does the lowest output price give the lowest bill?
No. Under the same-currency, listed-rate assumption, Groq and Together AI cost less than SambaNova when output tokens are fewer than seven times input tokens. SambaNova’s listed input price is 0.22 USD per million tokens against $0.15.1,2,3 The bills tie at seven output tokens per input token; above that ratio, SambaNova costs less. Cerebras’s quoted exploration-tier rates are higher than Groq’s in both directions.4
The break-even comes straight from the listed rates. SambaNova’s bill equals Groq’s when 0.07 × input tokens = 0.01 × output tokens, which is output at seven times input.
| Illustrative token mix | Groq | Together AI | SambaNova | Cerebras |
|---|---|---|---|---|
| 1M input + 1M output | $0.75 | $0.75 | $0.81 | $1.10 |
| 10M input + 1M output | $2.10 | $2.10 | $2.79 | $4.25 |
| 1M input + 10M output | $6.15 | $6.15 | $6.12 | $7.85 |
These mixes are computed from the listed prices. They aren’t observed workloads. The second row represents a longer input and shorter output, where SambaNova’s calculated total is roughly a third above Groq’s. In the third row, SambaNova saves three cents in about six dollars.
Use one real workload to test the claim. Count a week of input and output tokens, then apply each provider’s two prices. The ratio decides the order more than the output price does. What controls AI inference cost sets out the other levers, from caching to hardware.
Why rank on gpt-oss-120b instead of Llama 3.3 70B?
Of the two shared models tested, gpt-oss-120b has a printed per-token price on four of the five pages; Llama 3.3 70B has one on two. Groq lists Llama 3.3 70B as an Enterprise model with “ContactSales” in its price column.2 Cerebras’s Developer tier lists gpt-oss 120b and gemma-4-31b, not Llama.4 Fireworks AI’s landing page prints no model inference rates.5 That leaves SambaNova and Together AI as the two captured self-serve Llama 3.3 70B prices.
On that model the order flips. Together AI lists $1.04 per million tokens for both input and output.3 SambaNova lists 0.60 USD input and 1.20 USD output.1 A two-row ranking would put Together AI first, but it would leave Groq, the subject of the search, without a price.
Ranking one model keeps model size out of the comparison, so the remaining gap is the provider’s price. It doesn’t prove each provider serves identical weights at identical precision. None of the four ranked pages states the serving precision for gpt-oss-120b.
Which Groq competitors publish speed claims?
Groq and Cerebras print speed figures on the pages checked; SambaNova, Together AI and Fireworks AI print none. Groq lists 500 tokens per second for gpt-oss-120b.2 Cerebras lists “~3000 tokens/s” for GPT OSS 120B.4 Both are vendor claims. Groq’s table gives no test conditions. Cerebras’s footer says its performance comparisons rest on third-party benchmarking or internal testing and may vary by workload, configuration, date and model.4
On paper, Cerebras claims six times Groq’s speed for the same model, at a listed output price 25% higher. The pages can’t settle whether that holds. Neither discloses prompt length, concurrency or the point at which throughput was read, and no independent measurement sits in the evidence behind this ranking.
Groq prints speed for its other models in the same table: 280 tokens per second for Llama 3.3 70B and 1000 for gpt-oss-20b.2 Fireworks AI’s page names Standard, Priority and Fast serverless tiers and offers enterprise deployments “with faster speeds”, without a figure.5
The speed claims come from companies that design their own accelerators. Cerebras builds wafer-scale processors, SambaNova a reconfigurable dataflow unit and Groq its Language Processing Unit.7 Together AI and Fireworks AI price NVIDIA GPU capacity by the hour on their pricing pages, such as an NVIDIA HGX H100 at $3.99 an hour on Together AI’s GPU clusters3 and an H100 80 GB GPU at $8.00 an hour for Fireworks AI on-demand deployments.5 SambaNova designs its own chip but prints no speed on its pricing page.1
Which Groq competitors offer a free tier?
Two of the four competitors state free starting credit on their pricing pages. Cerebras offers “free $5 credit to start” on its self-serve Developer tier.4 Fireworks AI says “Get started with $1 in free credits.”5 SambaNova’s and Together AI’s pricing pages state no free credit, and nor does Groq’s models page. On output charges alone, $5 divided by Cerebras’s listed $0.75 yields about 6.7 million gpt-oss-120b output tokens; that illustration excludes input charges and does not establish credit eligibility or expiry.
Both providers publish paid usage billing alongside those starting credits. Cerebras describes its Developer tier as pay-as-you-go.4 Fireworks AI bills serverless use per token, postpaid.5 A starting credit is not evidence of a recurring free allowance.
Together AI lists one model, Ternary Bonsai 27B, at $0.00 per million tokens for input and output.3 That’s a zero-priced model, not account credit, and it isn’t gpt-oss-120b. One published list of Groq alternatives says Together AI, Fireworks AI and Baseten give credits before payment.8 Together AI’s own pricing page, checked 28 September 2026, doesn’t state that, so confirm it at sign-up. Groq’s models page shows rate limits under a “Developer plan” label without stating a free allowance.2
Why do other lists of Groq competitors disagree?
They define the competitor differently. CB Insights names chip companies, led by Hailo, Positron and Rebellions, as Groq’s top competitors.9 eesel’s list names model APIs: OpenAI, Anthropic, Together AI, Perplexity AI and Anyscale.8 Chipstrat, in October 2024, framed Groq as an Nvidia competitor that also competes with Google, AWS and Microsoft as cloud providers.10 This ranking answers the API buyer’s version: who sells the same open model by the token.
Prices drift as well. eesel’s FAQ quotes Groq’s Llama 3.1 8B Instant at $0.05 / $0.08.8 Groq’s models page, checked 28 September 2026, lists that model as Enterprise with “ContactSales” pricing.2 eesel’s Together AI figures for gpt-oss-120B, $0.15 / $0.60, match Together AI’s own page.3
Speed statements diverge the same way. eesel says Groq “still leads on raw tokens-per-second”.8 Cerebras’s own page prints ~3000 tokens/s for gpt-oss-120b against Groq’s printed 500.4,2 Neither statement is an independent measurement.
This is the first edition of the ranking, so there’s no earlier release to compare against.
What the ranking cannot tell you
The ranking covers one listed price for one model on one date. It doesn’t measure output quality, latency, uptime or rate limits under load, and it can’t confirm that each provider serves gpt-oss-120b at the same precision. It excludes negotiated enterprise rates, Groq’s Flex Processing and Batch Processing tiers2 and any provider whose pricing page omits the price.
It doesn’t carry across models either. On Qwen 3.8 27B, Cerebras lists $1.49 per million output tokens4 against $4.00 for Groq’s qwen/qwen3.8-27b, which Groq marks as a preview model.2 That’s the reverse of the gpt-oss-120b order.
The order changes when any ranked provider edits its gpt-oss-120b price row, or when Fireworks AI prints a rate on its pricing page. Confirm the current price on the provider’s page before committing spend. More pricing and serving research sits in the model infrastructure category, and new price checks go out in the Glamdring email briefing.
Frequently asked questions
What are some free alternatives to Groq?
Cerebras and Fireworks AI state free starting credit on their pricing pages checked 28 September 2026: $5 at Cerebras and $1 at Fireworks AI.4,5 Both publish paid usage billing, but those starting-credit statements do not establish a recurring free tier or all credit conditions. Together AI lists Ternary Bonsai 27B at $0.00 per million tokens, which is a zero-priced different model rather than free gpt-oss-120b credit.3
Is Cerebras faster than Groq?
Cerebras claims ~3000 tokens/s for gpt-oss-120b on its pricing page, and Groq claims 500 tokens per second for the same model.4,2 Both are vendor claims under conditions neither page fully discloses, so they don’t support a measured verdict. The claimed speed comes with a higher listed output price: $0.75 at Cerebras against $0.60 at Groq.
Can I still buy Llama 3.3 70B from Groq?
Groq lists llama-3.3-70b-versatile as an Enterprise model, with “ContactSales” in both its price and rate-limit columns, checked 28 September 2026.2 Self-serve per-token prices for Llama 3.3 70B appear at Together AI, $1.04 for input and output,3 and at SambaNova, 0.60 USD input and 1.20 USD output.1
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


