Together AI competitors: official prices, October 2026
On a USD comparison basis, SambaNova lists GPT OSS 120B output at 0.59 per million tokens; Groq and Fireworks tie at 0.60 and Cerebras lists 0.75. Confirm unlabelled billing currencies; Together AI's standard baseline remains unresolved.

On a USD comparison basis, SambaNova has the lowest listed standard output price among these four Together AI competitors for GPT OSS 120B: USD 0.59 per million output tokens, captured 28 September 2026.1 Groq and Fireworks list 0.60; Cerebras lists 0.75, with each source’s date retained below.2,3,4 Confirm unlabelled billing currencies before relying on the order. The finding covers one model’s output tariff, not the complete inference bill.
TL;DR
- Together AI’s standard GPT OSS 120B baseline remains unresolved, so the competitor order does not establish savings against Together.5,6
- Groq and SambaNova list free access to GPT OSS 120B with model-specific limits; Cerebras and Fireworks advertise starter credits.7,8,4,9
- GPU cluster prices buy time on hardware. Keep those rates separate from API token prices and compare the billing product as well as the number.5,9
How this ranking was built
We ordered four inference APIs by listed standard output price per million tokens for GPT OSS 120B, using each provider’s official model or pricing page.1,2,3,4 SambaNova and Fireworks explicitly identify USD; Groq and Cerebras print dollar amounts without a currency code.1,2,3,4 The model treats those dollar-only amounts as USD without conversion. Confirm billing currency in the quote; tax treatment is unspecified in the captured price tables. The comparison follows our research methodology.
Original price captures are dated 28 September 2026; supplementary official Fireworks serverless pricing, Together model confirmation, Groq Free Plan limits and SambaNova rate limits were captured on 1 October 2026.1,2,4,5,3,6,7,8 These are capture dates, not a claim that the providers released new prices on those days.1,2,3,4
The inclusion rule is a public tariff for the same GPT OSS 120B model and an identifiable non-Batch output price.1,2,3,4 Five providers were checked: four competitors enter the ranking and Together AI stays outside as the baseline whose billing mode needs clarification.1,2,3,4,5,6 Model-name variations are normalised to GPT OSS 120B; provider products remain separate.1,2,3,4
Equal prices share a rank. Input charges, cached-input charges, free access, credits, Batch discounts, enterprise quotes and GPU-hour prices do not determine that rank.1,2,3,4
Which Together AI competitors have the lowest output price?
SambaNova leads the GPT OSS 120B comparison at USD 0.59 per million output tokens; Groq and Fireworks tie at USD 0.60, followed by Cerebras at USD 0.75.1,2,3,4 The table orders reported list prices, with no performance weighting.1,2,3,4
| Rank | Inference API | GPT OSS 120B output (USD per million tokens) | Listed billing product | Checked date |
|---|---|---|---|---|
| 1 | SambaNova | 0.59 | Published API tariff | 28 September 20261 |
| 2= | Groq | 0.60 | Production-model API tariff | 28 September 20262 |
| 2= | Fireworks | 0.60 | Standard serverless | 1 October 20263 |
| 4 | Cerebras | 0.75 | Developer pay-as-you-go | 28 September 20264 |
SambaNova is the first option to test if the decision is the lowest listed output tariff for this model.1,2,3,4 Groq and Fireworks have the same output rate in this comparison, so that field alone cannot choose between them.2,3 Cerebras distinguishes exploration workloads from its quote-based Enterprise tier; its listed Developer price should not be read as an enterprise-capacity quote.4
Fireworks’ GPT OSS 120B Standard cell is input/cached input/output, with USD 0.60 output, not a Batch rate.3 Its documentation explicitly separates Standard from Priority and states that Batch inference is billed at 50% of serverless input and output pricing.3 Reading the first number in the cell as the output price, or applying the Batch discount to a standard request, would change the comparison’s meaning.3
Where does Together AI fit in the comparison?
Together AI lists GPT OSS 120B, but its standard output-price baseline remains outside the competitor ranking because the captured pricing material does not resolve the billing mode.5,6 The Together model page confirms availability, not a missing rate or mode.6
The 28 September pricing capture places a GPT OSS 120B row beneath both “Price per 1M tokens” and “Batch API price”.5 The 1 October model page identifies the endpoint openai/gpt-oss-120b and lists serverless and dedicated deployment options.6 Those details confirm that Together offers the shared model; they do not settle which billing mode the earlier price row represents.5,6
Confirm the standard tariff for the endpoint before calculating a saving against Together AI. A model-availability confirmation and a clearly labelled billing rate answer different procurement questions.5,6
Which competitors offer free access?
Groq and SambaNova publish free-plan access to GPT OSS 120B with model-specific limits, while Cerebras and Fireworks advertise starter credits.7,8,4,9 Compare the free allowance separately from the paid output tariff.1,2,7,8
| Provider | Free access or starting credit | Conditions in the captured source | Checked date |
|---|---|---|---|
| Groq | Free Plan includes openai/gpt-oss-120b | 30 requests/minute; 1,000 requests/day; 8,000 tokens/minute; 200,000 tokens/day | 1 October 20267 |
| SambaNova | Free Tier includes gpt-oss-120b | No payment method linked with the account; 20 requests/minute; 20 requests/day; 200,000 tokens/day | 1 October 20268 |
| Cerebras | USD 5 starting credit | Developer tier is self-serve pay-as-you-go | 28 September 20264 |
| Fireworks | USD 1 starting credit | Serverless uses per-token pricing and postpaid billing | 28 September 20269 |
SambaNova’s no-payment-method condition defines its Free Tier: linking a payment method puts the account in the Developer Tier.8 Its GPT OSS 120B free row allows 20 requests per day, so the daily request ceiling can bind before the token allowance does.8
Groq applies rate limits at organisation level and says an account can hit whichever limit comes first.7 Its documentation describes the published rows as a high-level summary with possible exceptions and directs users to their account’s exact limits.7 Neither provider’s token-per-day allowance is stated as an output-only allowance.7,8
The Cerebras and Fireworks pages describe finite starting credits, which should be budgeted separately from continuing free access.4,9 Their captured pricing pages do not specify expiry or payment-method requirements for those credits.4,9 Confirm those terms during signup.
How do GPU cluster prices compare?
Together AI publishes on-demand GPU Cluster rates, while Fireworks publishes on-demand deployment rates; both charge for GPU time.5,9 Those hourly rates belong outside the API token-price ranking.5,9 The following prices were captured on 28 September 2026, with tax treatment unspecified.5,9
| GPU | Together AI GPU Clusters, on-demand (USD per GPU per hour) | Fireworks on-demand deployments (USD per GPU per hour) |
|---|---|---|
| H100 | 3.995 | 8.009 |
| H200 | 5.995 | 8.009 |
| B200 | 8.195 | 13.009 |
| B300 | 9.995 | 15.009 |
Together labels all cluster prices per GPU per hour and publishes separate preemptible, on-demand and reserved categories.5 Fireworks states that its on-demand deployments bill per GPU second and gives hourly equivalents; region-restricted deployments carry a 1.5x premium.9 The matching hourly denominator does not make a cluster and a managed deployment identical products.5,9
Keep Together’s GPU Clusters rate card separate from its Dedicated Inference table, which carried an H100 promotion valid until 30 September 2026.5 The cluster table above uses the listed on-demand cluster rate, without extending that separate promotion.5
For a wider hardware comparison, use the H100 rental-price comparison and B200 rental-price comparison. To compare hardware time with an API bill, first establish how many billable tokens your workload delivers during the paid GPU time. The price pages alone do not supply that workload result.1,2,3,4,5,9
What the ranking cannot tell you
Glamdring Research’s price ordering establishes the lowest listed standard output tariff for GPT OSS 120B within this captured competitor set.1,2,3,4 It does not establish the lowest total bill, a performance winner or a saving against Together AI’s unresolved standard baseline.1,2,3,4,5,6
Each API price page separates input from output charges.1,2,3,4 A complete cost comparison needs your input volume as well as your generated output, and any applicable cached-input treatment.1,2,3,4 Equal output prices therefore leave part of the cost decision open.2,3
Use the same prompts and output limits to test candidate APIs. Record billed usage, completed responses and response times under your intended load. Check the rate limits your account receives before treating free access as production capacity.7,8
The next official price or rate-limit change should trigger a refresh of this comparison. A clearly identified Together standard tariff would also allow a direct baseline comparison.5,6 Continue with our model-infrastructure research or the published research archive.
Frequently asked questions
Are there cheaper Batch rates?
Fireworks states that Batch inference costs 50% of its serverless input and output pricing in the documentation captured on 1 October 2026.3 That discount belongs to Batch billing and is excluded from the standard-output ranking.3 Compare Batch offers separately if your workload can use that mode.
Does an open-weight model mean the API is free?
An open-weight model and free hosted API access are separate things: Together lists GPT OSS 120B as an open model with deployment options, while Groq publishes both a paid model tariff and Free Plan limits.6,2,7 SambaNova also separates Free and Developer access according to whether a payment method is linked.8 Check the hosting provider’s plan for the model you intend to call.
Will adding a payment method change SambaNova access?
Yes. SambaNova’s documentation captured on 1 October 2026 applies the Developer Tier when a payment method is linked with the account.8 For GPT OSS 120B, the Developer row lists 60 requests per minute and 12,000 requests per day; Developer accounts also have a 20 million-token daily limit across all models.8 The larger request allowance remains subject to that account-wide token ceiling.8
Sources checked
Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.


