GPT-Realtime 1.5
OpenAI's flagship audio-in, audio-out model for voice agents and customer support workflows.
Pricing
Input$4 / 1M text, $32 / 1M audio
Cached input$0.4 / 1M text, $0.4 / 1M audio
Output$16 / 1M text, $64 / 1M audio
Image input is priced separately at $5 / 1M.
Cost bandPremium
Free tierNo free API tier.
Limits
Context window32k
Max output4,096
Knowledge cutoffSep 30, 2024
DeploymentOpenAI Realtime API
OpenAI-compatible APIYes
ModalitiesText, Image input, Audio input and output
Where it fits
Best forProduction voice assistants, live support, phone and headset interfaces.
Business fitUse when speech quality and low-turn latency matter more than text-token economics.
Capabilities
RealtimeVoice agentsStreaming
Benchmark scores
Composite71 / 100
Coding65 / 100 (HumanEval (vendor))
Reasoning68 / 100 (MMLU-Pro (vendor))
Long context60 / 100
Vision62 / 100
Instruction following78 / 100 (MT-Bench (vendor))
Editorial aggregate of published benchmark results as of May 1, 2025. These are not independent measurements run by this site.
Change history
No recorded changes for this model yet. Pricing and context are checked daily; changes appear here with dates and sources.
Data
Last checkedMar 31, 2026
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.