Black Forest Labs announced the launch of FLUX3, a multi-modal foundation model, in Early Access mode. The model adopts a unified architecture for joint learning of images, videos, and audio. Based on the Self-Flow (self-supervised flow matching) learning framework, it synchronously trains videos, images, and audio, further expanding from the FLUX.1 and FLUX.2 series to multi-modal generation and understanding tasks.

Completely superior performance, native audio synchronization output
FLUX3 can output videos up to 20 seconds long in a single generation, along with native audio. It supports various creation modes such as text-to-video, image-to-video, video-to-video, input video and audio continuation, keyframe-to-video, multilingual dialogue, and multi-shot video concatenation. In manual evaluation of 10-second 720p videos with sound, FLUX3 achieved a win rate of 69% against Grok Imagine Video, and 52% against Seedance 2.0 and Gemini Omni Flash, demonstrating a comprehensive leading advantage.
Comprehensive image capabilities, entering robot behavior prediction
