Google Launches Gemini 3.8 Live Voice Models
Real-time voice agents have typically relied on separate speech recognition, language-model and text-to-speech layers, adding latency and making tool-heavy tasks feel fragmented. Google’s Gemini 3.8 Live lineup uses native audio-to-audio models designed to keep dialogue flowing while handling visual context, reasoning and external functions. The launch matters because it pushes conversational AI beyond demos toward customer service, booking and enterprise workflows where response speed, continuity and reliable task completion are critical.
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on Sept. 15, 2026, making both available through the Gemini API and Google AI Studio. Extended Thinking can reason and call asynchronous tools while continuing to speak. Google said it scored 82.6 on Artificial Analysis’ Speech to Speech Quality Index, 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. Standard API pricing is $0.005 per minute for audio input and $0.018 per minute for audio output.
All Coverage
2 original reportsThe Backstory
The history behind this eventGoogle Launches Gemini 3.5 Transcribe to Filter Fillers and Noise
Google is expanding its Gemini audio lineup for transcription, meeting notes, captions, customer service and voice-controlled applications. Speech-to-text systems often lose accuracy when speakers overlap, conversations are interrupted or background noise obscures words. Gemini 3.5 Transcribe is designed to use context to recognize specialized terminology and produce cleaner prose, reducing the manual editing typically required after automated transcription.
The new model supports more than 85 languages and recorded an average word error rate of 2.6%, according to Google. It can distinguish as many as three speakers, remove fillers such as “ums” and “ahs,” and provide real-time translation while handling noisy or interrupted speech. Traditional Chinese is not currently listed among the supported languages, limiting its immediate usefulness for some Taiwan-based users.
Google Launches Three Gemini Models for AI Agents
The generative AI race is shifting from chatbots toward agents that can call tools and complete multi-step tasks, making inference speed, token consumption and cybersecurity increasingly important to enterprise buyers. Google is expanding the lightweight Flash tier to lower the cost of running such workloads at scale, while strengthening the Gemini API ecosystem as competition intensifies among model providers.
As of July 29, 2026, Google has unveiled three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. The releases emphasize faster computation and greater token efficiency for agentic workloads, while the cybersecurity-tuned Flash Cyber has entered a pilot focused on vulnerability remediation. Google said Gemini 3.5 Pro remains in testing, leaving a gap in the lineup, and that pretraining for the next-generation Gemini 4 has begun.
Google Launches Gemini 3.5 Live Translate for Real-Time Speech Interpretation in More Than 70 Languages
Google is extending Gemini 3.5's generative AI capabilities to real-time interpretation through direct speech-to-speech translation that preserves a speaker's original tone, cadence and delivery. Compared with conventional systems that first convert speech into text before synthesizing audio, the new model can reduce pauses, making it significant for international meetings, travel and customer service communications.
Google has launched the Gemini 3.5 Live Translate audio model, which can translate more than 70 languages in real time. The model is being integrated into Google Translate and Google Meet and is also available in developer preview. Google has begun testing it with partners including Southeast Asian ride-hailing and delivery platform Grab to assess translation speed and naturalness in real-world service settings.
Google Launches Gemini 3.1 Flash TTS With Support for 70 Languages and Scene Direction
Google is extending Gemini’s multimodal capabilities into speech generation, enabling developers to create more natural interactions for customer service, video dubbing and accessibility services. Gemini 3.1 Flash TTS supports multiple languages and emotional control, lowering the barrier to producing voice content across markets and making it particularly important for localizing AI applications.
As of July 19, 2026, Google AI developer relations lead Logan Kilpatrick announced the launch of Gemini 3.1 Flash TTS. The model supports more than 70 languages and lets developers use audio tags to direct scenes and adjust emotion and delivery. It is now available to try in Google AI Studio and can be integrated through the Gemini API, though pricing has not yet been announced.
Google Launches Gemini 3.1 Flash Live, Expands Search Live to More Than 200 Countries
Google first launched Search Live in the United States in September 2025, extending generative AI beyond text search to continuous voice- and camera-based follow-up questions. The service combines Google’s search index with AI Mode, breaking down questions during conversations and providing links to webpages. The move reflects Google’s effort to reshape its core search gateway with Gemini. The company did not disclose the investment behind the latest rollout.
On March 26, 2026, Google released Gemini 3.1 Flash Live, a real-time audio model designed to improve the speed, naturalness and stability of multilingual responses. It also expanded Search Live to more than 200 countries and territories where AI Mode is available, with support for 98 languages. Users can tap Live in the Google app on Android or iOS, or use Google Lens to conduct continuous searches by combining the camera with voice queries.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →