Google Tests Gemini 3.8 Flash as 3.5 Pro Faces Third Delay
Google is accelerating development of its Gemini artificial-intelligence models as it competes with OpenAI and Anthropic. The Flash line is designed to balance reasoning performance, speed and deployment costs, while the flagship Pro tier is expected to handle more demanding tasks. Delays to the higher-end model therefore carry broader implications for Google’s ability to narrow the gap with rivals and translate research advances into widely available products.
Google has begun internal testing of its next-generation Gemini 3.8 Flash model, signaling that development of the faster product line remains on a rapid cadence. By contrast, Gemini 3.5 Pro has reportedly been postponed three times amid an underlying architectural overhaul and persistent reasoning bottlenecks. The company had not set a firm release date as of Aug. 31, 2026, leaving it primarily reliant on Flash models to sustain its competitive position.
All Coverage
1 original reportsThe Backstory
The history behind this eventGoogle Unveils Gemini 3.7 Flash in August AI Push
Google is tightening the links between its Gemini models, Pixel hardware and productivity services as it seeks to turn advances in artificial intelligence into widely used products. Its August rollout emphasized lower-cost tools for developers, voice-driven workflows and on-device assistance, showing how the company is shifting from headline model performance toward practical applications across coding, communications, media creation and consumer devices.
In a Sept. 1, 2026 roundup, Google said Gemini 3.7 Flash arrived just three weeks after 3.6 Flash, with introductory pricing at half the earlier model’s original cost per million tokens. The company also introduced Gemini 3.5 Transcribe and the Pixel 11 lineup, powered by the Tensor G6 chip and the latest Gemini Nano. The Gemini app surpassed 1 billion monthly users, while Gemini Omni 1.1 Flash added scene extension, frame interpolation and 4K upscaling for video production.
Google Launches Gemini 3.7 Flash With Half-Price API Offer
Google’s Flash lineup is positioned as a workhorse tier that balances speed, cost and reasoning capability for high-volume coding and AI-agent workloads. Rapid model cycles and falling inference costs are becoming central to competition among Google, OpenAI and other providers seeking to lock developers and enterprises into their platforms. Gemini 3.7 Flash first surfaced in leak reports and a briefly visible Python SDK pull request, signaling that another release was imminent.
Google formally released Gemini 3.7 Flash on Aug. 13, 2026, just over three weeks after unveiling Gemini 3.6 Flash on July 21. Standard Gemini API pricing is set at an introductory $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, half the regular rates. Prices rise to $1.50 and $7.50, respectively, on Jan. 1, 2027. Google describes the model as its most capable Flash offering for agentic workflows and multimodal reasoning.
Google Rolls Out Faster, Cheaper Gemini Models
Google is positioning Gemini as the foundation for AI agents that can plan, call tools and complete multi-step tasks across software and, increasingly, robotics. Its Flash family targets the middle ground between frontier-grade intelligence and production economics, where latency, token use and reliability determine whether companies can deploy agents at scale. The strategy also links Google DeepMind’s models with consumer gateways including Search and the Gemini app, intensifying competition for developers and everyday AI users.
On July 21, Google DeepMind released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The company said 3.6 Flash uses 17% fewer output tokens than 3.5 Flash and costs $1.50 per million input tokens and $7.50 per million output tokens. Flash-Lite generates 350 output tokens a second and is priced at $0.30 for input and $2.50 for output. Both are available through the Gemini API and Google AI Studio, while Flash-Lite is rolling out in Google Search.
Google Launches Three Gemini Models for AI Agents
The generative AI race is shifting from chatbots toward agents that can call tools and complete multi-step tasks, making inference speed, token consumption and cybersecurity increasingly important to enterprise buyers. Google is expanding the lightweight Flash tier to lower the cost of running such workloads at scale, while strengthening the Gemini API ecosystem as competition intensifies among model providers.
As of July 29, 2026, Google has unveiled three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. The releases emphasize faster computation and greater token efficiency for agentic workloads, while the cybersecurity-tuned Flash Cyber has entered a pilot focused on vulnerability remediation. Google said Gemini 3.5 Pro remains in testing, leaving a gap in the lineup, and that pretraining for the next-generation Gemini 4 has begun.
Google Delays Gemini 3.5 Pro as AI Integration Strains Mount
Gemini 3.5 Pro is intended to be Google’s next flagship artificial-intelligence model as the company races OpenAI and Meta in increasingly competitive foundation models. The project matters because Google is betting that its strengths in multimodal processing and access to Search data can differentiate Gemini, while tighter coordination across research, infrastructure and consumer products is critical to turning those technical advantages into commercially viable services.
The model’s launch has reportedly slipped by several months after Google opted to rebuild it from scratch, compounding coding bottlenecks, internal bureaucracy and factional disputes. The difficulties have fueled employee frustration and talent departures, while an internal evaluation was reported to show the model potentially trailing Meta’s entry-level Muse Spark. Alphabet shares fell 4.4% following one delay report. Google has since launched an internal effort called “Antigravity” to consolidate research resources and accelerate development.
Google Unveils Gemini 3.5 Flash for Autonomous Agents and Software Development
Google introduced Gemini 3.5 Flash at its 2026 Google I/O developer conference, extending its lineup of multimodal Gemini models and targeting AI applications that require frequent calls, long context windows and real-time responses. The model delivers performance exceeding that of the previous-generation Pro at Flash-tier cost while strengthening autonomous-agent and software-development capabilities.
The newly announced Gemini 3.5 Flash outperformed the previous-generation Pro in several agent benchmarks, with reasoning and processing speeds improving fourfold. Google also demonstrated that the model could build an operating system within 12 hours at a total cost of less than $1,000. The company reiterated plans to increase investment in AI infrastructure, saying it was seeing real market demand.
Google Launches Gemini 3.1 Flash-Lite for Low-Cost, Large-Scale Workloads
Google DeepMind’s Gemini 3 lineup comprises Pro, Flash and Flash-Lite, designed respectively for advanced reasoning, speed and cost efficiency. Flash-Lite targets high-volume tasks such as translation, content moderation and data labeling, lowering the barrier for businesses deploying generative AI at scale.
Google released a preview of Gemini 3.1 Flash-Lite on March 3, 2026, making it available through Google AI Studio and Vertex AI. The API charges $0.25 per 1 million input tokens and $1.50 per 1 million output tokens, about one-eighth the price of the Pro version. Google said output speed increased by 45%.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →