On this page3 sections
Black Forest Labs released FLUX 3, a multimodal foundation model that generates images, videos with audio, and predicts physical actions within a single unified architecture.
The model represents a departure from training separate systems for each modality. Instead, FLUX 3 learns simultaneously from images, videos, and audio to build what the company calls "a representation of the world" — understanding how objects hold together, how things move, and how events sound.
Black Forest Labs built FLUX 3 on Self-Flow, its approach for aligning multimodal generation and understanding within the same architecture. The company scaled up compute and data resources to train across video, images, and audio simultaneously.
Video generation leads capabilities
FLUX 3 can create videos with native audio up to 20 seconds long. Core capabilities include text-to-video generation, image-to-video animation, video-to-video generation using reference clips, and keyframe-to-video generation for controlled transitions.
The model handles multilingual dialogue and generates content across visual styles from candid camcorder footage to animation and cinematics. It can chain individual clips into longer sequences lasting several minutes.
In preliminary evaluations, FLUX 3 was preferred over competing models in head-to-head comparisons: 77% preference over Runway Gen-4.5, 69% over Grok Imagine Video, and 93% over Luma Ray 3.2.
FLUX 3 Video is available in early access, with the company noting these are preliminary results expected to improve during development.
Physical AI integration
The model extends beyond content creation to action prediction. Black Forest Labs partnered with mimic robotics to develop FLUX-mimic, combining the FLUX 3 backbone with specialized robot learning for dexterous manipulation tasks.
This integration is being tested on production tasks at Audi, demonstrating the company's thesis that physical AI and content creation share the same foundational requirements.
Rollout timeline
Black Forest Labs plans to release capabilities over the coming months through early access phases. The roadmap includes video and audio generation through APIs, action prediction for research and commercial partners, image synthesis and editing, and eventual open-weight access to the multimodal backbone.
FLUX 3 Image will enter early access in the following weeks, with the company conducting preliminary evaluations showing significant improvements in complex prompt handling and multilingual text generation compared to earlier FLUX versions.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.