Black Forest Labs has opened early access to FLUX 3, a visual intelligence model that generates video, images, and synchronised audio. The German company best known for its AI image work is now also deploying the same technology on the factory floor through a partnership with Audi.
What FLUX 3 can do
FLUX 3 produces videos up to 20 seconds long with native audio. The system supports multilingual dialogue, custom typography, and consistent style or character adherence across clips. It accepts text, image, and video as inputs, and Black Forest Labs says internal testing shows its outputs outperform rivals such as Runway, Kling, and Grok Imagine.
The full rollout will include video, image, and partner-specific action models. An open-weight FLUX 3 Dev variant is also planned, sized to run on factory hardware.
From pixels to robots
Through FLUX-mimic, built with Zurich-based mimic robotics, the same visual intelligence learns factory tasks from roughly half an hour of demonstration data. That compares favourably to the 30-plus hours typical of existing approaches.
Why it matters
Black Forest Labs has traditionally moved slower than some competitors, but this unified release marks a clear acceleration. Longer, audio-synced videos with coherent characters show the payoff. FLUX-mimic also signals that visual models are steadily entering the physical world, becoming a core component in the robotics stack.


