Google Releases Experimental DiffusionGemma Model, Quadrupling GPU Text-Generation Speed
Google's Gemma is a family of open models available to developers. Traditional large language models typically generate text autoregressively, predicting one token at a time. DiffusionGemma instead uses diffusion-based generation to process multiple tokens simultaneously, with a focus on improving GPU utilization and reducing text-generation latency.
As of July 20, 2026, Google had released the experimental DiffusionGemma model. Built on the Gemma 4 family and incorporating a mixture-of-experts architecture, or MoE, it can process 256 tokens in parallel. On dedicated GPUs, it generates text at up to four times the speed of conventional models, though its output quality still trails Gemma 4.
All Coverage
2 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →