GPT-Realtime 1.5

OpenAI · GPT Realtime · voice · Current stable voice model · gpt-realtime-1.5

OpenAI's flagship audio-in, audio-out model for voice agents and customer support workflows.

Pricing

Input$4 / 1M text, $32 / 1M audio
Cached input$0.4 / 1M text, $0.4 / 1M audio
Output$16 / 1M text, $64 / 1M audio

Image input is priced separately at $5 / 1M.

Cost bandPremium
Free tierNo free API tier.

Limits

Context window32k
Max output4,096
Knowledge cutoffSep 30, 2024
DeploymentOpenAI Realtime API
OpenAI-compatible APIYes
ModalitiesText, Image input, Audio input and output

Where it fits

Best forProduction voice assistants, live support, phone and headset interfaces.
Business fitUse when speech quality and low-turn latency matter more than text-token economics.

Capabilities

RealtimeVoice agentsStreaming

Benchmark scores

Composite71 / 100
Coding65 / 100 (HumanEval (vendor))
Reasoning68 / 100 (MMLU-Pro (vendor))
Long context60 / 100
Vision62 / 100
Instruction following78 / 100 (MT-Bench (vendor))

Editorial aggregate of published benchmark results as of May 1, 2025. These are not independent measurements run by this site.

Change history

No recorded changes for this model yet. Pricing and context are checked daily; changes appear here with dates and sources.

Data

Last checkedMar 31, 2026

Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.

Open in the comparison table JSON API