Google Launches Gemini 3.7 Flash as OpenAI Previews Ultrafast GPT-5.6 Sol
Low latency and inference cost are becoming as important as raw model capability as companies deploy AI agents, which can make repeated model calls while completing multistep work. Google has positioned its Flash line for high-volume, cost-sensitive tasks, while OpenAI’s GPT-5.6 Sol targets complex professional workloads. The latest releases sharpen competition over how quickly advanced systems can respond and how economically developers can operate them at scale.
On Aug. 13, 2026, Google released Gemini 3.7 Flash across its developer and enterprise channels, pricing it through Dec. 31 at $0.75 per million input tokens and $3.75 per million output tokens. OpenAI the same day previewed GPT-5.6 Sol Ultrafast, a Cerebras-powered service tier capable of generating as many as 750 output tokens per second, or up to 14 times the speed of Standard processing. Access is initially limited to selected API customers, and OpenAI has not disclosed pricing.
All Coverage
1 original reportsThe Backstory
The history behind this eventGoogle Reportedly Launches Gemini 3.7 Flash With API Price Cuts
Google’s Gemini Flash family is designed for high-volume inference, coding and agentic workloads where latency and cost matter as much as benchmark performance. A cheaper Gemini 3.7 Flash would intensify competition with OpenAI and Anthropic, while lowering the cost of deploying AI agents and software-development tools at scale. The reported pricing shift also underscores how model providers are increasingly using inference costs as a competitive lever.
As of Aug. 14, 2026, reports said Google had released Gemini 3.7 Flash at $0.75 per 1 million input tokens, with both input and output charges reportedly cut to about half their previous levels through year-end. A pull request briefly added the model to Google’s official Python SDK on GitHub before being renamed and closed roughly 15 seconds later. Google’s definitive pricing documentation and full availability details remained to be confirmed.
Google Rolls Out Faster, Cheaper Gemini Models
Google is positioning Gemini as the foundation for AI agents that can plan, call tools and complete multi-step tasks across software and, increasingly, robotics. Its Flash family targets the middle ground between frontier-grade intelligence and production economics, where latency, token use and reliability determine whether companies can deploy agents at scale. The strategy also links Google DeepMind’s models with consumer gateways including Search and the Gemini app, intensifying competition for developers and everyday AI users.
On July 21, Google DeepMind released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The company said 3.6 Flash uses 17% fewer output tokens than 3.5 Flash and costs $1.50 per million input tokens and $7.50 per million output tokens. Flash-Lite generates 350 output tokens a second and is priced at $0.30 for input and $2.50 for output. Both are available through the Gemini API and Google AI Studio, while Flash-Lite is rolling out in Google Search.
Google Launches Three Gemini Models for AI Agents
The generative AI race is shifting from chatbots toward agents that can call tools and complete multi-step tasks, making inference speed, token consumption and cybersecurity increasingly important to enterprise buyers. Google is expanding the lightweight Flash tier to lower the cost of running such workloads at scale, while strengthening the Gemini API ecosystem as competition intensifies among model providers.
As of July 29, 2026, Google has unveiled three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. The releases emphasize faster computation and greater token efficiency for agentic workloads, while the cybersecurity-tuned Flash Cyber has entered a pilot focused on vulnerability remediation. Google said Gemini 3.5 Pro remains in testing, leaving a gap in the lineup, and that pretraining for the next-generation Gemini 4 has begun.
Google Unveils Gemini 3.5 Flash for Autonomous Agents and Software Development
Google introduced Gemini 3.5 Flash at its 2026 Google I/O developer conference, extending its lineup of multimodal Gemini models and targeting AI applications that require frequent calls, long context windows and real-time responses. The model delivers performance exceeding that of the previous-generation Pro at Flash-tier cost while strengthening autonomous-agent and software-development capabilities.
The newly announced Gemini 3.5 Flash outperformed the previous-generation Pro in several agent benchmarks, with reasoning and processing speeds improving fourfold. Google also demonstrated that the model could build an operating system within 12 hours at a total cost of less than $1,000. The company reiterated plans to increase investment in AI infrastructure, saying it was seeing real market demand.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.