Mark RadarMARK RADAR
About
EN
Sign in
Event File AI Google AI Models

Google Releases Experimental DiffusionGemma Model, Quadrupling GPU Text-Generation Speed

2 reports · First detected 2026-06-11 · Last active 2026-06-11

Google's Gemma is a family of open models available to developers. Traditional large language models typically generate text autoregressively, predicting one token at a time. DiffusionGemma instead uses diffusion-based generation to process multiple tokens simultaneously, with a focus on improving GPU utilization and reducing text-generation latency.

As of July 20, 2026, Google had released the experimental DiffusionGemma model. Built on the Gemma 4 family and incorporating a mixture-of-experts architecture, or MoE, it can process 256 tokens in parallel. On dedicated GPUs, it generates text at up to four times the speed of conventional models, though its output quality still trails Gemma 4.

All Coverage

2 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)