Back to feed
News Story
meng shao (X)
3 sources

FLUX 3 Released: Black Forest Labs Unveils 'Real World Model'

Black Forest Labs has released FLUX 3, positioning it as a 'real world model' rather than just a text-to-video or text-to-image model. The model uses a multimodal flow architecture that simultaneously processes images, video, audio, and language to learn how the world works, supporting both content generation and embodied action prediction. Video features are now in early access, with image and action capabilities coming soon.

SynthePulse Insight · AI deep reading

Flux 3 Released: Native Audio-Video Generation and Open-Source Roadmap

Version 1 · 2 sources

Black Forest Labs releases Flux 3, achieving native audio generation for video for the first time, and plans to open-source model weights.

  • Flux 3 supports video generation up to 20 seconds and adds native audio for the first time.
  • In BFL internal tests, Flux 3 leads most competitors in 10-second 720p video preference rate, but with limited advantage.
  • The model is based on the Self-Flow method, unifying image, video, audio, and motion.
  • BFL plans a phased release, ultimately open-sourcing the Flux 3 Dev model.
  • Flux-mimic video motion model has been tested in Audi production tasks.
Open section navigationCore Capabilities: Native Audio and Multimodal Unification

Core Capabilities: Native Audio and Multimodal Unification

Black Forest Labs released Flux 3 on July 23, 2026, a multimodal foundation model that simultaneously learns images, video, and audio. Flux 3 is the first to support generating videos with native audio, up to 20 seconds. BFL believes that a single modality cannot fully capture reality, and joint training of the three can fill information gaps.

Flux 3 supports text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue, and smart editing (stitching clips into multi-shot sequences). BFL states the model excels at matching human facial expressions and sounds with physical events.

The model is based on the Self-Flow method, using a multimodal Transformer that converts images, video, audio, and motion into a shared internal representation via dedicated encoders and decoders, then converts back to output. The motion component provides a foundation for robotics applications.

Performance Comparison: Internal Tests Lead but with Limited Advantage

BFL conducted early evaluations using 10-second 720p videos for user preference tests. Flux 3 achieved a 93% preference rate over Luma Ray 3.2, 77% over Runway Gen-4.5, and 69% over Grok Imagine Video.

The advantage narrows against stronger competitors: 60% over Kling v3 Pro, 59% over Happy Horse v1, 57% over Happy Horse 1.1, and 52% over both Seedance 2.0 and Gemini Omni Flash. BFL notes the results are preliminary and no independent tests have been conducted.

Release Plan and Open-Source Commitment

BFL plans a phased release: Flux 3 Video is already available, Flux 3 Image will enter early access within weeks, and motion prediction features are initially limited to partners.

BFL will release an open-source weights version, Flux 3 Dev, but currently it is only accessible via early access application. The API will open in the coming weeks, followed by the open-source version.

Robotics Application: Flux-mimic and Audi Testing

BFL collaborated with Mimic Robotics to develop Flux-mimic, a video motion model that has been tested in Audi production tasks. This demonstrates the potential of Flux 3's unified learning framework in robotics.

Credibility boundary

This article is primarily based on BFL's official blog and company statements. Performance data are internal test results and have not been independently verified. The open-source plan is solely announced by BFL, and actual timelines may change.

Insight takeaway

Flux 3 integrates native audio into video generation for the first time and plans to open-source, but its performance advantage is limited and lacks independent verification.