Google Launches Gemini 3.1 Flash-Lite for Low-Cost, Large-Scale Workloads
Google DeepMind’s Gemini 3 lineup comprises Pro, Flash and Flash-Lite, designed respectively for advanced reasoning, speed and cost efficiency. Flash-Lite targets high-volume tasks such as translation, content moderation and data labeling, lowering the barrier for businesses deploying generative AI at scale.
Google released a preview of Gemini 3.1 Flash-Lite on March 3, 2026, making it available through Google AI Studio and Vertex AI. The API charges $0.25 per 1 million input tokens and $1.50 per 1 million output tokens, about one-eighth the price of the Pro version. Google said output speed increased by 45%.
All Coverage
3 original reportsThe Backstory
The history behind this eventGoogle Unveils Gemini 3.5 Flash for Autonomous Agents and Software Development
Google introduced Gemini 3.5 Flash at its 2026 Google I/O developer conference, extending its lineup of multimodal Gemini models and targeting AI applications that require frequent calls, long context windows and real-time responses. The model delivers performance exceeding that of the previous-generation Pro at Flash-tier cost while strengthening autonomous-agent and software-development capabilities.
The newly announced Gemini 3.5 Flash outperformed the previous-generation Pro in several agent benchmarks, with reasoning and processing speeds improving fourfold. Google also demonstrated that the model could build an operating system within 12 hours at a total cost of less than $1,000. The company reiterated plans to increase investment in AI infrastructure, saying it was seeing real market demand.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.