Google Launches Three Gemini Models for Faster, Cheaper AI Agents
Enterprise AI agents can consume large volumes of tokens as they reason, call tools and complete multistep tasks, making inference costs and latency critical barriers to deployment. Google is expanding its lighter Gemini Flash tier to address those constraints and strengthen its developer ecosystem as it competes with OpenAI and Anthropic. The lineup also extends into cybersecurity, an area where specialized models can help identify and repair software vulnerabilities.
Google on July 22, 2026, unveiled three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. The releases promise faster processing and lower token costs, with Flash Cyber entering a pilot program for cybersecurity tasks including vulnerability remediation. A larger Gemini 3.5 Pro model remains in testing, leaving a gap in the current lineup, while Google said pretraining for the next-generation Gemini 4 has begun.
All Coverage
8 original reportsThe Backstory
The history behind this eventGoogle Unveils Lower-Cost Gemini Cybersecurity Model
As AI agents become faster at finding software flaws than defenders can patch them, the economics of repeated code scanning are becoming a central cybersecurity concern. Google DeepMind built Gemini 3.5 Flash Cyber on its lightweight Gemini 3.5 Flash model to discover, validate and repair vulnerabilities efficiently, positioning it as a lower-cost alternative to large security systems such as Anthropic’s Mythos. The smaller design lets CodeMender invoke the model repeatedly and search more code paths without relying on one costly frontier-model call.
Google DeepMind unveiled the model on July 21, 2026, and said it would soon be available exclusively to governments and trusted partners through a limited-access CodeMender pilot; no standalone price was disclosed. In a fixed-invocation test on the V8 JavaScript engine, Flash Cyber found 55 confirmed unique issues, versus 47 for Gemini 3.5 Flash and 36 for Claude Opus 4.6, including 10 missed by both. Google also said the model uncovered remote-code-execution and memory-corruption flaws in two hours during internal work.
Google Releases Gemma 4 Open Models to Bolster Its Position in Open AI Market
Google positions the Gemma family as a line of downloadable, self-deployable open models, offering a range of parameter sizes to lower barriers to generative AI adoption for businesses and developers. Gemma 4 takes aim at Meta’s Llama and Alibaba’s Qwen as the companies compete in local inference, tool integration and developer ecosystems.
Google’s newly released Gemma 4 comes in four configurations ranging from 2B to 31B parameters, supports more than 140 languages and multimodal processing, and is optimized for NVIDIA H100 GPUs. The 12B version can run locally on a consumer laptop with 16GB of VRAM and is licensed under Apache 2.0. Available information does not specify an exact release date or any commercial value.
Google Unveils Gemini 3.5 Flash for Autonomous Agents and Software Development
Google introduced Gemini 3.5 Flash at its 2026 Google I/O developer conference, extending its lineup of multimodal Gemini models and targeting AI applications that require frequent calls, long context windows and real-time responses. The model delivers performance exceeding that of the previous-generation Pro at Flash-tier cost while strengthening autonomous-agent and software-development capabilities.
The newly announced Gemini 3.5 Flash outperformed the previous-generation Pro in several agent benchmarks, with reasoning and processing speeds improving fourfold. Google also demonstrated that the model could build an operating system within 12 hours at a total cost of less than $1,000. The company reiterated plans to increase investment in AI infrastructure, saying it was seeing real market demand.
Google Launches Gemini 3.1 Pro, Expands Model to Cloud and Enterprise Platforms
After launching Gemini 3 Pro on November 18, 2025, Google continued extending generative AI from consumer applications into enterprise workflows. Gemini 3.1 Pro focuses on complex, multistep reasoning and can integrate spreadsheets with unstructured data. For Google, the model is a key part of its effort to compete for the enterprise AI infrastructure market through Vertex AI and Gemini Enterprise.
Google released a preview of Gemini 3.1 Pro on February 19, 2026, and introduced it to Vertex AI and Gemini Enterprise. The model scored 77.1% on ARC-AGI-2, more than twice the score of 3 Pro. JetBrains measured improvements of up to 15%, while Databricks said its OfficeQA performance led the industry. Google did not disclose licensing fees.
Google Launches Gemini 3.1 Flash-Lite for Low-Cost, Large-Scale Workloads
Google DeepMind’s Gemini 3 lineup comprises Pro, Flash and Flash-Lite, designed respectively for advanced reasoning, speed and cost efficiency. Flash-Lite targets high-volume tasks such as translation, content moderation and data labeling, lowering the barrier for businesses deploying generative AI at scale.
Google released a preview of Gemini 3.1 Flash-Lite on March 3, 2026, making it available through Google AI Studio and Vertex AI. The API charges $0.25 per 1 million input tokens and $1.50 per 1 million output tokens, about one-eighth the price of the Pro version. Google said output speed increased by 45%.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.