EPISODE · Jul 27, 2026 · 7 MIN
Black Forest Labs FLUX 3: from video generation to robot control
from Air Street Press
Description:Black Forest Labs has launched FLUX 3, a multimodal foundation model that learns jointly from images, video and audio within a single architecture, and mimic robotics has released FLUX-mimic, a video-action model built on that backbone and being tested on real assembly work in Audi's Production Lab. Nathan Benaich of Air Street Capital, an investor in Black Forest Labs, reads his Air Street Press essay on why a model trained to predict how scenes evolve turns out to be a usable robot controller. Covers the FLUX 3 preference results against Runway Gen-4.5, Grok Imagine Video, Kling v3 Pro, Seedance 2.0 and Gemini Omni Flash; how a lightweight action decoder reads the video prediction path without ever generating video; the frozen-backbone ablation against π0.5; the 101-millisecond system reaction time on a single RTX 5090; and what Air Street's own robotics deal flow says about where the constraint really sits.Chapters (estimated at ~150 wpm, slide proportionally against final audio):0:00 A robot arm in Audi's Production Lab0:40 What Black Forest Labs and mimic released1:20 Early access and open weights1:50 The preference-test results2:40 What a model must represent to predict video3:30 Reading actions off the video path4:20 The frozen-backbone ablation5:00 101 milliseconds, and the Audi tasks5:50 What we see in robotics deal flow6:30 The road to physical intelligenceLinks: FLUX 3 · FLUX-mimic (BFL) · FLUX-mimic (mimic) · Odyssey Series B · BFL Series B ·
Embed this episode
NOW PLAYING
Black Forest Labs FLUX 3: from video generation to robot control
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.