Glamdring

Model infrastructure

OpenAI Realtime API pricing: OpenAI price list, Sep 2026

The two mini models list the lowest realtime audio prices: $10 per million input tokens and $20 per million output tokens. Full-size models charge $32 and $64. These are OpenAI's prices checked 28 September 2026.

Glamdring Research14 min6 sources checked

Cutaway engraving of a conference microphone and loudspeaker connected to an audio and network appliance

OpenAI’s gpt-realtime-2.1 costs $32.00 per million audio input tokens, $0.40 per million cached audio input tokens and $64.00 per million audio output tokens. Its text tokens cost $4.00 in, $0.40 cached and $24.00 out. The mini version, gpt-realtime-2.1-mini, charges $10.00, $0.30 and $20.00 for the same audio lines and $0.60, $0.06 and $2.40 for text. These are the rates from OpenAI’s pricing pages, checked 28 September 2026.1,2

Audio output is the dearest line on every token-billed realtime model OpenAI lists, at twice the audio input rate.3

TL;DR

  • The two mini models have the lowest listed audio rates: $10 per million input tokens and $20 per million output tokens. Full-size models cost 3.2 times as much on those lines. Voice quality needs a separate workload test.3
  • Cached input is the deepest discount on the price list. Cached audio input on gpt-realtime-2.1 costs 1.25% of the fresh rate; on the mini it costs 3%. Cached text and image input cost 10% of the fresh rate on every realtime model.3
  • gpt-realtime-2 and gpt-realtime-2.1 carry identical prices. The older gpt-realtime and gpt-realtime-1.5 differ only on text output, at $16.00 per million against $24.00.3
  • OpenAI’s live voice, translation and transcription models bill by the minute: gpt-live-1 at $0.05, gpt-realtime-translate at $0.034 and live transcription at $0.017.3
  • OpenAI’s pricing pages give no audio tokens-per-minute figure. A per-minute cost for the token-billed models has to come from your own usage data.1,3

How these prices were checked

Every price here is read from OpenAI’s own pricing pages, captured on 28 September 2026 (UTC+10) as part of Glamdring’s model infrastructure research.1,3 We captured openai.com/api/pricing and the developer pricing table at 15:05. We captured the developer table again at 21:12 with every “All models” list expanded, which shows the four older realtime models.

The field is OpenAI’s Realtime and audio generation models table. It gives Input, Cached input and Output / cost columns for each modality, in prices per 1M tokens unless noted.3

The inclusion rule is every model in that table whose ID begins with gpt-realtime. The expanded table lists 12 models and six qualify. The other six are gpt-audio-1.5, gpt-audio-mini, gpt-audio, gpt-4o-mini-tts, tts-1 and tts-1-hd. None of them has a cached-input price, and tts-1 and tts-1-hd are priced per million characters.3 Four time-billed realtime products sit in their own section below. Nothing was merged.

The openai.com page lists only gpt-realtime-2.1 and its mini. Its 16 figures for those two models match the developer table exactly.1,2 Both pages show prices with a $ sign. Neither states a currency code or tax treatment in the captured text. Our research methodology sets out how Glamdring captures and dates vendor prices.

What does each realtime model cost per million tokens?

Four full-size realtime models charge $32.00 per million audio input tokens and $64.00 per million audio output tokens, and the two mini models charge $10.00 and $20.00.3 The full-size models split on text output: $24.00 per million on gpt-realtime-2.1 and gpt-realtime-2, and $16.00 on gpt-realtime-1.5 and gpt-realtime.3 The table below gives all 48 per-token rates for the six models in OpenAI’s own order.

ModelAudio in ($/1M)Audio cached in ($/1M)Audio out ($/1M)Text in ($/1M)Text cached in ($/1M)Text out ($/1M)Image in ($/1M)Image cached in ($/1M)
gpt-realtime-2.1$32.00$0.40$64.00$4.00$0.40$24.00$5.00$0.50
gpt-realtime-2.1-mini$10.00$0.30$20.00$0.60$0.06$2.40$0.80$0.08
gpt-realtime-2$32.00$0.40$64.00$4.00$0.40$24.00$5.00$0.50
gpt-realtime-1.5$32.00$0.40$64.00$4.00$0.40$16.00$5.00$0.50
gpt-realtime-mini$10.00$0.30$20.00$0.60$0.06$2.40$0.80$0.08
gpt-realtime$32.00$0.40$64.00$4.00$0.40$16.00$5.00$0.50

Source: OpenAI developer pricing table, checked 28 September 2026.3

Image tokens are input only. OpenAI gives no image output price for any realtime model.3 OpenAI describes gpt-realtime-2.1 as its “most capable model for realtime voice interactions” and the mini as its “strongest mini model yet”.1 Those are OpenAI’s descriptions. The price list is the part a buyer can check.

What do audio tokens cost on the Realtime API?

Audio tokens on the OpenAI Realtime API cost $32.00 per million in and $64.00 per million out on the four full-size models, and $10.00 in and $20.00 out on the two mini models.3 Cached audio input costs $0.40 per million on the full-size models and $0.30 on the minis.3 Per 1,000 tokens, gpt-realtime-2.1 charges $0.032 for audio input and $0.064 for audio output.

Audio output costs twice audio input on all six models. The mini discount is the same on both lines: $32.00 against $10.00 in, $64.00 against $20.00 out, a factor of 3.2 each time. The 3.2 multiple applies to fresh audio input and audio output; cached audio input has a smaller gap. The quality case for the full-size model has to cover the price difference in your own tests.

The constraint for budgeting is the unit. OpenAI prices realtime audio in tokens, and the captured pricing pages give no conversion from tokens to minutes of speech.1,3 A per-minute or per-call figure for these models therefore depends on token counts from your own sessions. For the wider cost drivers behind model serving, see what controls AI inference cost.

What do text and image tokens cost on the Realtime API?

Text tokens on the Realtime API cost $4.00 per million in and $24.00 out on gpt-realtime-2.1 and gpt-realtime-2, $4.00 and $16.00 on gpt-realtime-1.5 and gpt-realtime, and $0.60 and $2.40 on the two mini models.3 Image input costs $5.00 per million on the full-size models and $0.80 on the minis.3 Cached text and image input cost one tenth of those rates.

Realtime text is dearer than text on OpenAI’s general models. gpt-6-sol charges $2.00 per million input tokens and $10.00 per million output tokens at standard short-context rates.2 gpt-realtime-2.1 charges twice that input rate and 2.4 times that output rate. The gap is wider at the small end: gpt-6-luna charges $0.10 in and $0.50 out, against $0.60 and $2.40 on gpt-realtime-2.1-mini.2,3

The decision changes when a voice product also handles typed input. Text sent through a realtime session pays the realtime text rate. In October 2024 one developer in OpenAI’s community forum described adding a text input panel that bypasses the Realtime API to keep costs down.4 Routing typed turns to a text model is a design choice with a price attached, and the price list shows the size of it.

Images follow their own rate. OpenAI says text models price image tokens at standard text token rates, while GPT Image and gpt-realtime use a separate image token rate.1 On gpt-realtime-2.1 that rate is $5.00 per million, above its $4.00 text input rate.3

How much does cached input save?

Cached input is the deepest discount in OpenAI’s realtime price list. Cached audio input on gpt-realtime-2.1 costs $0.40 per million tokens against $32.00 fresh, 98.75% less.3 On gpt-realtime-2.1-mini the cached audio rate is $0.30 against $10.00, 97% less. Cached text and image input cost 10% of the fresh rate on all six realtime models.3

Caching applies to input only. No realtime model has a cached output rate, so audio output at $64.00 or $20.00 per million is untouched by it.3

The cache also narrows the gap between model sizes. Fresh audio input on gpt-realtime-2.1 costs 3.2 times the mini rate. Cached audio input costs $0.40 against $0.30, or 1.33 times. Cached audio on gpt-realtime-2.1 even costs less per token than fresh text input on the mini, at $0.40 against $0.60.3

Check how your usage reports cached tokens before you rely on the discount. When OpenAI added cached pricing to the Realtime API in October 2024, a developer in the announcement thread found cached-token details in the usage field of response.done events but not in the Playground’s usage total.5 That report is from 2024. Confirm the current behaviour in your own usage data.

Do older realtime models cost less?

Older full-size realtime models cost less only on text output. gpt-realtime and gpt-realtime-1.5 charge $16.00 per million text output tokens, while gpt-realtime-2 and gpt-realtime-2.1 charge $24.00, which is 50% more.3 Every audio, text input and image rate is identical across the four full-size models. The two mini models share one price set.3

Moving to an older full-size model saves a third of the text output cost and nothing on audio. For a voice product where most tokens are audio, the saving is small.

The comparison sits inside one price list captured on one day. It is a comparison across model versions, and we hold no earlier capture to show how any single model’s price moved over time. OpenAI’s current openai.com pricing page names only gpt-realtime-2.1 and its mini. The four older models appear in the developer table’s expanded list.1,3 The captured pages give no retirement date for them.

Which OpenAI realtime models are billed per minute?

OpenAI bills four realtime voice products by time: gpt-live-1 at $0.05 per minute, gpt-realtime-translate at $0.034 per minute, and gpt-live-transcribe and gpt-realtime-whisper at $0.017 per minute.3 An hour costs $3.00, $2.04 and $1.02 respectively.

ModelUsePrice per minutePrice per hour (computed)
gpt-live-1Voice sessions$0.05$3.00
gpt-realtime-translateLive translation$0.034$2.04
gpt-live-transcribeLive transcription$0.017$1.02
gpt-realtime-whisperLive transcription$0.017$1.02

Source: OpenAI developer pricing table, checked 28 September 2026; hourly figures are the per-minute price times 60.3

The gpt-live-1 hourly figure is a floor. OpenAI bills gpt-live-1 sessions per second, without rounding up to a whole minute, and charges backend model and tool usage separately.3 The openai.com page gives the per-second rates as $0.00083 for gpt-live-1, $0.00057 for gpt-realtime-translate and $0.00028 for gpt-live-transcribe.1 gpt-transcribe, at $0.0045 per minute, is for asynchronous and batch work, so it sits outside live use.1

These minute prices can’t be lined up against the token-billed models without a tokens-per-minute rate, and OpenAI’s pricing pages don’t give one.

How do you estimate a Realtime API bill?

Glamdring Research’s reading of the price list is that a Realtime API bill turns on two numbers: the share of audio input billed at the cached rate, and the volume of audio output. Audio output carries the highest rate on every realtime model, and cached audio input carries the deepest discount.3 The illustration below uses invented token counts to show the arithmetic. It is not a measured session.

Assume one session uses 100,000 audio input tokens, of which 60,000 bill at the cached rate, and 20,000 audio output tokens.

LineTokensgpt-realtime-2.1gpt-realtime-2.1-mini
Fresh audio input40,000$1.28$0.40
Cached audio input60,000$0.024$0.018
Audio output20,000$1.28$0.40
Total120,000$2.58$0.82
Total with no caching120,000$4.48$1.40

Rates from OpenAI’s developer pricing table, checked 28 September 2026.3

In this example caching cuts each bill by about 42%. Audio output then makes up about half of the gpt-realtime-2.1 bill, and no cache rate touches it.

Use one real session to test the estimate. OpenAI’s usage dashboard shows tokens used in the current and past billing cycles.1 A monthly budget stops requests once it is reached, but OpenAI says there may be a delay in enforcing the limit and the customer is responsible for any overage.1

Azure OpenAI meters realtime usage on its own terms. A Microsoft moderator answered in July 2025 that Azure bills from the portal metrics processed_prompt_tokens and generated_completion_tokens, and that the Realtime API’s own token counts are not what Azure uses for cost.6

Voice activity detection also moves the bill. In October 2024 a developer in OpenAI’s community forum reported that background noise can trigger a generation, and described a push-to-talk button that made costs more predictable.4

What these prices cannot tell you

OpenAI’s price list fixes the token and minute rates. It leaves open how many tokens a minute of speech produces, which discounts and uplifts reach realtime models, and how the models perform.

The captured pricing pages give no audio tokens-per-minute rate, so per-minute or per-call costs for token-billed models need your own usage data.1,3

The developer table gives Batch, Flex and Fast mode rates for flagship models but lists one rate set for realtime models.3 OpenAI charges a 10% uplift on regional processing endpoints for eligible models released on or after March 5, 2026. The captured text doesn’t say whether realtime models are eligible.3 FedRAMP endpoints carry a 10% uplift over the corresponding standard model rates.3 Prices appear with a $ sign and no stated currency code or tax treatment.

The price list says nothing about voice quality, reasoning or response speed. Test those on your own workload. Azure OpenAI publishes its own pricing page and billing metrics, so an Azure deployment needs its own check.6

New pricing checks across model infrastructure go out in Glamdring’s email briefing.

Frequently asked questions

How much do 1,000 tokens cost on the OpenAI Realtime API?

On gpt-realtime-2.1, 1,000 audio input tokens cost $0.032 and 1,000 audio output tokens cost $0.064. Text costs $0.004 per 1,000 input tokens and $0.024 per 1,000 output tokens. On gpt-realtime-2.1-mini the same four lines cost $0.01, $0.02, $0.0006 and $0.0024. Each figure is OpenAI’s per-million price divided by 1,000.3

Is the OpenAI Realtime API free or paid?

The OpenAI Realtime API is paid per token or per minute. OpenAI bills Playground usage the same as regular API usage. API access is billed separately from ChatGPT Plus, Business, Enterprise and Edu subscriptions, so a ChatGPT plan doesn’t cover Realtime API calls.1

How much does the rest of the OpenAI API cost?

OpenAI’s flagship text models range from $0.10 per million input tokens on gpt-6-luna to $50.00 per million output tokens on gpt-6-astra, at standard short-context rates. gpt-6-sol sits between them at $2.00 in and $10.00 out.2 For other providers’ rates on similar work, see our comparison of OpenAI API alternatives.

Does the Batch API discount apply to realtime models?

OpenAI’s captured pricing pages don’t publish a Batch rate for realtime models. The Batch API saves 50% on inputs and outputs and runs tasks asynchronously over 24 hours.1 The developer table lists Batch rates for flagship models and one rate set for realtime models.3 A 24-hour asynchronous window also doesn’t fit a live voice conversation. Confirm in the current quote before you budget on a discount.

Sources checked

  1. Business Pricing | OpenAIChecked September 28, 2026
  2. Realtime API extremely expensiveChecked September 28, 2026

Company-owned pages establish what a company says. They do not prove a market conclusion. Each source is dated so readers can judge each claim.