OpenAI Launches GPT-Realtime-2.1 Voice Model
Voice agents used in customer service, sales and assistant applications need fast responses, accurate recognition and reliable handling of interruptions. Excessive latency makes conversations feel unnatural. OpenAI has therefore built GPT-Realtime-2.1 on GPT-Realtime-2, aiming to lower the technical barriers to deploying real-time voice products at scale.
As of July 20, 2026, OpenAI has released GPT-Realtime-2.1 with improved recognition of letters and numbers, silence and noise handling, and interruption behavior. Reports said latency was at least 25% lower than in the previous generation. The model supports a 128,000-token context window and up to 32,000 output tokens. Audio input and output cost $32 and $64 per million tokens, respectively.
All Coverage
1 original reportsThe Backstory
The history behind this eventOpenAI Launches GPT-Realtime-2 Voice Model With GPT-5-Level Reasoning
Real-time voice agents must do more than respond quickly. They need to understand intent as conversations unfold, retain context, handle interruptions and corrections, and reliably call tools to complete tasks. By bringing GPT-5-level reasoning directly into a voice model, OpenAI aims to narrow the gap between natural conversation and real-world execution, bringing customer service, travel and cross-language applications closer to deployable working interfaces.
OpenAI released GPT-Realtime-2 on May 7, 2026, alongside Realtime-Translate and Realtime-Whisper. The main model's context window increased from 32K to 128K, with five selectable reasoning-effort levels. Translation supports more than 70 input languages and 13 output languages, while audio is priced at $32 per million input tokens and $64 per million output tokens.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.