Glamdring

Model infrastructure

Together AI pricing: GPU clusters and inference, Sep 2026

On Together AI's pricing page, checked 28 September 2026, the HGX H100 is the cheapest cluster GPU at $3.99 per GPU-hour on demand, falling to $3.19 on a 91-to-180-day reservation and $1.99 preemptible; the B300 tops the list at $9.99.

Glamdring Research13 min4 sources checked

Engraving of a cluster of GPU servers joined by thick bundles of high-speed interconnect cables

The NVIDIA HGX H100 is the cheapest GPU in Together AI’s clusters at $3.99 per GPU-hour on demand, on Together AI’s pricing page checked 28 September 2026.1 A 91-to-180-day reservation cuts it to $3.19, and preemptible capacity costs $1.99.1 Serverless inference is billed per unit of work, mostly per million tokens, with no minimums.2 One dated promotion is live: dedicated H100 inference at $3.99 an hour instead of $5.49, valid until 30 September 2026.1

TL;DR

  • Longer cluster terms cut the price unevenly. The H200 drops 33.4% on a 91-to-180-day reservation. The B200 drops 17.1% on the same term and only 2.4% on 7 to 30 days.1
  • Preemptible capacity costs almost exactly half of on-demand on all four priced GPUs, from $1.99 (H100) to $4.99 (B300).1
  • The GB200 NVL72, the GB300 NVL72 and every term of 181 days or longer need a sales conversation. So do all B300 reservations.1
  • Provisioned throughput costs $0.05 per unit per minute. At full use it matches the serverless list rate per token for all three listed models, so a unit buys reserved capacity at the serverless token price.1
  • After the promotion ends, a dedicated H100 endpoint costs $1.50 an hour more than an on-demand cluster H100.1
  • Together AI’s own docs and pricing page disagree on some cached-input rates and billing units. Glamdring uses the pricing page.3,1

How this ranking was built

Glamdring Research ranked the hardware in Together AI’s GPU Clusters table by one field: the ON-Demand price, which the page states per GPU per hour.1 It sits in our model infrastructure category and covers Together AI’s published GPU cluster and inference prices, checked 28 September 2026. The page was captured at 15:05 Australian Eastern Standard Time and shows no last-updated date.

The GPU Clusters table lists six hardware rows. Four carry an on-demand price and are ranked, cheapest first. The GB200 NVL72 and GB300 NVL72 show a dash for on-demand and preemptible and “Contact us” for reserved, so they sit unranked.1 The captured table prints two header rows, which leaves the column order ambiguous. We checked it against the page’s separate on-demand table. All four on-demand prices match, so the order is preemptible, on-demand, then the four reserved terms.1

Prices carry a $ sign. The page names no currency code and no tax treatment.1 Discounts are Glamdring’s arithmetic: one minus the term price divided by the on-demand price, rounded to one decimal place. Monthly figures assume a 720-hour month of continuous use. Our research methodology sets out how prices are captured and checked.

What does Together AI charge per GPU-hour for a cluster?

Together AI’s on-demand cluster prices run from $3.99 per GPU-hour for the HGX H100 to $9.99 for the HGX B300, on the pricing page checked 28 September 2026.1 Every figure below is per GPU per hour, as the page states it.

RankGPUOn-demand ($ per GPU-hour)PreemptibleReserved 7–30 daysReserved 31–90 daysReserved 91–180 daysReserved 181+ days
1NVIDIA HGX H100$3.99$1.99$3.69$3.45$3.19Contact us
2NVIDIA HGX H200$5.99$2.99$4.99$4.15$3.99Contact us
3NVIDIA HGX B200$8.19$4.09$7.99$7.79$6.79Contact us
4NVIDIA HGX B300$9.99$4.99Contact usContact usContact usContact us
UnrankedNVIDIA GB200 NVL72Not listedNot listedContact usContact usContact usContact us
UnrankedNVIDIA GB300 NVL72Not listedNot listedContact usContact usContact usContact us

Price rises in rank order: H100, H200, B200, then B300.1 One crossover matters for buyers. A 91-to-180-day H200 reservation costs $3.99, the same as an on-demand H100. At that term, an H200 costs no more than pay-as-you-go H100 capacity.1

The page labels the cheapest column “Preemptible Compute” and “Pay as you go”. The page doesn’t describe the preemption terms. Read the service terms before you put a training run on it.1 A banner on the same page announces on-demand B200s on Together GPU Clusters.1 Shared filesystem storage, colocated with the compute, is listed at $0.16 per GiB per month.1

How much do reserved terms cut Together AI’s cluster prices?

Reserved terms cut Together AI’s cluster prices by 2.4% to 33.4% off on-demand, depending on the GPU and the term. Preemptible capacity cuts about 50.1% on every priced GPU.1

GPU7–30 days31–90 days91–180 daysPreemptible
NVIDIA HGX H1007.5%13.5%20.1%50.1%
NVIDIA HGX H20016.7%30.7%33.4%50.1%
NVIDIA HGX B2002.4%4.9%17.1%50.1%
NVIDIA HGX B300n/an/an/a50.1%

The H200 has the steepest term curve. The B200 hardly moves until the 91-day threshold: $8.19 on demand, $7.99 for 7 to 30 days, $7.79 for 31 to 90 days, then $6.79.1 A B200 buyer with a 60-day job saves 40 cents an hour against on-demand by reserving. A buyer who can stretch to 91 days saves $1.40.

In cash terms, one H100 for a 720-hour month costs $2,872.80 on demand and $2,296.80 on a 91-to-180-day reservation. The H200 falls from $4,312.80 to $2,872.80. The B200 falls from $5,896.80 to $4,888.80.1 The estimate assumes continuous use. The captured page doesn’t say whether a reservation bills for the whole term, how many GPUs a cluster must hold, or what the 181-day-plus rates are. Confirm all three in the current quote.

What promotions is Together AI running, and when do they end?

Together AI’s pricing page shows one promotion with an end date. Dedicated inference on an NVIDIA HGX H100 costs $3.99 per GPU-hour instead of $5.49, and the page says it’s valid until 09/30/26: 30 September 2026.1

That’s a 27.3% cut. It applies to Together AI’s dedicated inference product; the GPU Clusters table carries no promotion. For a 720-hour month, the promotional H100 endpoint costs $2,872.80 against $3,952.80 at the standard price.1 The promotion runs two days past our check date. Any contract signed after 30 September 2026 should be priced at $5.49 unless Together AI states otherwise in writing.

No other row on the captured page carries a promotional label or an end date.1 Two standing discounts are not promotions. Cached input tokens are billed at a lower rate on some chat models. Batch workloads are discounted by up to 50% on some models.2

How does Together AI price serverless inference?

Together AI prices serverless inference per unit of work, with no minimums and no provisioning cost. Chat, language, embedding and rerank models bill per input and output token.2 The pricing page quotes chat rates per million tokens, image models per image or per megapixel, text-to-speech per million characters, transcription per audio minute and video per video.1

Among the paid chat models on the page, input runs from $0.09 per million tokens (Qwen3.8 Flash) to $3.00 (Kimi K3). Output runs from $0.14 (Llama 3 8B Instruct Lite) to $15.00 (Kimi K3).1 Llama 3.3 70B is $1.04 each way, and gpt-oss-120B is $0.15 in and $0.60 out.1

Cached input is where the list price moves most. Fourteen of the 27 models in the first chat table show a cached rate. DeepSeek V4.1 Flash drops from $0.30 to $0.006 per million input tokens, a 98% cut. Qwen3.5-397B-A17B drops from $0.60 to $0.35, a 41.7% cut.1 Caching is automatic and prefix-based. It’s also best-effort on serverless: the cache is shared across the fleet, and hits aren’t guaranteed. Dedicated endpoints enable prompt caching by default, scoped to your own replicas.2 Budget on the standard input rate and treat cached savings as upside, unless your traffic runs on a dedicated endpoint.

Image pricing adds a steps factor. The docs give cost as megapixels × price per megapixel × (steps ÷ default steps). Steps count only above the default.3 The page prices FLUX.2 [max] at $0.070 per megapixel with 50 default steps, and GPT Image 2 at $0.053 per image.1 Our explainer on what controls AI inference cost covers the other inputs to a serving bill.

Is Together AI’s provisioned throughput cheaper than serverless?

No. Glamdring Research’s arithmetic shows that a fully used provisioned throughput unit (PTU) costs the same per token as Together AI’s serverless list rate, for all three models priced on 28 September 2026.1 Each PTU costs $0.05 per minute and delivers a fixed tokens-per-minute rate that varies by model and token type.1

ModelToken typeTokens per minute per PTU$ per 1M tokens at full useServerless list price ($ per 1M)
MiniMax M3Input166,667$0.30$0.30
MiniMax M3Cached input833,333$0.06$0.06
MiniMax M3Output41,667$1.20$1.20
Kimi K3Input16,667$3.00$3.00
Kimi K3Cached input166,667$0.30$0.30
Kimi K3Output3,333$15.00$15.00
GLM-5.2Input35,714$1.40$1.40
GLM-5.2Cached input192,308$0.26$0.26
GLM-5.2Output11,364$4.40$4.40

The full-use figure is $0.05 divided by the tokens per minute, multiplied by one million, rounded to the cent. Every one of the nine rows lands on the serverless price.1 A PTU running at half its capacity costs twice the serverless rate per token.

The page states its savings estimate against “the selected commercial model’s published list price”, not against Together AI’s own serverless rate. It assumes continuous provisioning of about 43,800 minutes a month.1 At $0.05 a minute, that is $2,190 per PTU per month. The calculator’s default of 353 PTUs matches its displayed $773,070 monthly estimate.1

The case for PTUs is capacity. Serverless models are rate-limited, and Together AI’s docs point steady traffic that needs higher limits or reserved hardware to provisioned throughput or dedicated endpoints.2 Buy PTUs when throughput is the constraint. At full use the token price stays the same, and every idle minute raises it.

What does Together AI’s dedicated inference cost compared with a cluster GPU?

Dedicated inference costs more per GPU-hour than the same GPU in a cluster. The HGX H100 is $5.49 against $3.99, and the HGX B200 is $8.99 against $8.19, on the pricing page checked 28 September 2026.1

The H100 premium is $1.50 an hour, or 37.6%. The B200 premium is $0.80, or 9.8%.1 Until 30 September 2026, the promotional H100 dedicated price of $3.99 removes the H100 premium entirely. The H200, B300, GB200 NVL72 and GB300 NVL72 have no public dedicated price, and every reserved dedicated term routes to sales.1

The premium buys a managed endpoint. The page describes single-tenant GPU instances with guaranteed performance, custom-model support and autoscaling.1 A dedicated endpoint also lets you pin hardware to a region. Serverless requests run in a region you can’t select.2

Why do other published Together AI prices disagree?

Other published Together AI prices disagree for three traceable reasons: Together AI’s docs differ from its own pricing page, the docs describe units two ways, and a third-party tracker carries models the pricing page’s serverless tables don’t list.

The docs model catalogue lists Qwen3.8-2.4T-A95B cached input at $0.50 and Qwen3.7 Max cached input at $0.50 per million tokens.3 The pricing page shows $0.25 and $0.30.1 The docs list Qwen3.8 Flash output at $0.282.3 The pricing page shows $0.28.1

The docs overview says video bills per second of output, and speech-to-text and text-to-speech per second of audio.2 The docs catalogue and the pricing page price video per video, transcription per audio minute and speech per million characters.3,1 Check the invoice unit on a small test job before modelling a media workload.

AI Pricing Guru’s Together AI page is titled July 2026 but says it was last updated on 27 September 2026. It lists DeepSeek V4 Pro at $1.74 in and $3.48 out, Kimi K2.6 at $1.20 and $4.50, and GPT-OSS 20B at $0.05 and $0.20.4 The captured pricing page lists DeepSeek V4 Pro 0813 at $1.32 and $3.96. Kimi K2.6 and gpt-oss-20B appear only in its fine-tuning tables, not its serverless chat tables.1 Where figures differ, this page uses Together AI’s pricing page.

What the ranking cannot tell you

These are Together AI’s listed prices on 28 September 2026. They are a company statement. They don’t show what reserved or enterprise customers negotiate.1

The page shows price. Nor does it show available capacity, delivery times, minimum cluster size or preemption terms. The ranking doesn’t measure performance per dollar. The captured page states no currency code or tax treatment; confirm both in the current quote. The page offers a “Batch API price” view, but the capture doesn’t show which figures belong to it. The batch discount here comes from the docs’ “up to 50%”.2

The ranking will be rebuilt when Together AI changes its pricing page. The H100 promotion’s 30 September 2026 end date is the first scheduled change. New pricing research goes out in our free research briefing.

Frequently asked questions

Is Together AI free to use?

Mostly not. One chat model, Ternary Bonsai 27B, is listed at $0.00 for input and output on the pricing page checked 28 September 2026. Everything else on the page is metered.1 Serverless use carries no minimum charge, so you pay only for what you process.2 The captured pricing page states no free-credit amount.

How much do 1,000 tokens cost on Together AI?

Divide the per-million rate by 1,000. Llama 3.3 70B at $1.04 per million tokens costs $0.00104 per 1,000 tokens, input or output.1 Across the paid chat models on the page, 1,000 input tokens cost $0.00009 to $0.003. The same volume of output costs $0.00014 to $0.015, the top figure being Kimi K3.1

What are the disadvantages of Together AI’s serverless pricing?

Serverless trades control for simplicity. Together AI’s docs say serverless models are rate-limited and you can’t select the serving region. Cached-input discounts are best-effort, with no guaranteed hit rate or retention window.2 Each fix costs money: provisioned throughput or a dedicated endpoint. Provisioned throughput matches the serverless token rate only at full use.1

Sources checked

  1. PricingChecked September 28, 2026
  2. Overview - Together AI docsChecked September 28, 2026
  3. Available models - Together AI docsChecked September 28, 2026

Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.