Black Forest Labs Unveils FLUX 3 for Video and Robotics
Black Forest Labs, known for the FLUX.1 and FLUX.2 image-generation models, is pushing beyond still pictures with FLUX 3, a unified multimodal foundation model trained across images, video and audio. The model also supports action prediction, linking generative media with robotics through a shared representation of motion, contact and cause and effect. That approach positions the same underlying AI system to create audiovisual content and help machines manipulate objects in the physical world, broadening competition in both creative AI and industrial automation.
Black Forest Labs opened FLUX 3 to Early Access on July 23, 2026. The model can produce videos of up to 20 seconds with native synchronized audio in a single generation, using text, image or video inputs. The company also developed FLUX-mimic with mimic robotics, adapting FLUX 3’s video backbone for robot action prediction; the system has been tested on production tasks at Audi. Black Forest Labs said video prediction consumed more than 95% of total training compute, while pricing and parameter count were not disclosed.
All Coverage
3 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →