Google Unveils Gemini Omni Multimodal Video Model With Conversational Editing
Google unveiled Gemini Omni at I/O 2026, touting its ability to combine text, images, video and other inputs to generate high-fidelity video. The model advances video production beyond one-off prompts by enabling repeated adjustments through natural-language “conversational editing,” lowering barriers to professional editing and content creation.
The latest version, Gemini Omni Flash, can precisely modify camera angles, lighting and motion. The initial release caps video output at 10 seconds. Google said the model achieved state-of-the-art results in three video-generation benchmarks. The Gemini Omni API is coming soon, while free integration with YouTube Shorts is expected during the week of I/O 2026.
All Coverage
4 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.