Black Forest Labs Unveils FLUX 3 AI Image and Video Models

Black Forest Labs has launched FLUX 3 with support for 20-second video generation with synchronized audio.

TL;DR
  • FLUX 3 Capabilities: FLUX 3 generates images and videos up to 20 seconds with synchronized sound from one multimodal model.
  • Access Limits: Only approved users can test Video and Action; broader software access and downloadable model parameters remain planned.
  • Benchmark Limits: Black Forest Labs’ preliminary 720p comparisons lack published sample sizes, full methodology, and independent evaluation.
  • Robot Testing: FLUX-mimic converts video features into robot commands, and Audi tested it on production and logistics tasks.

This week, Black Forest Labs made FLUX 3 available in Early Access. FLUX 3 Video can generate up to 20 seconds of video with synchronized sound, while a related Action variant extends the system into robotics. Only approved early users can test either variant.

FLUX 3 is not yet fully available and remains in early access. Prospective users must request access rather than use the public API. Pricing, service commitments, downloadable model parameters, full benchmark methodology, image benchmarks, and the resolution available at the 20-second limit also remain undisclosed so far.

FLUX 3 is one underlying AI system trained across images, video, and audio. Dedicated encoders turn each medium into a shared internal representation, and decoders translate it back into media or actions.

Joint training is designed to generate sound with movement rather than pass audio to a separate model. Shared representations can also feed robot planning and control workflows.

One Model, Several Output Modes

FLUX 3 Video accepts text, image, and video inputs. It supports continuation and keyframe transitions, multilingual dialogue, and multi-shot chaining, but users must still request early access. At launch, native ComfyUI support was reported absent while access remains gated.

Creators can use the model family to prompt a new clip, modify existing footage, continue its picture and sound, or connect several shots. Multiple generation modes can create a unified image-and-video production workflow with fewer handoffs among editing, storyboarding, video variation, and localization. Access controls leave those production options untested by outside teams.

 

Agent-controlled chaining can join generated shots, while keyframes give creators explicit transition points between segments. Continuation spans picture and sound, so a team could extend an existing clip without rebuilding the audio track in a separate tool.

Black Forest Labs Flux 3 - Self-Flow

Company-run preliminary preference comparisons provide an initial performance signal based on 10-second, 720p clips. Both FLUX 3 and the test harness remained in development.

FLUX 3 led Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77%. Results against Seedance 2.0 and Gemini Omni Flash were 52% each.

Missing sample sizes, rater counts, detailed methodology, image benchmarks, and independent evaluation keep the company-run results preliminary. Tests on shorter clips cannot establish quality or resolution across the full 20-second limit, much less a reproducible performance lead.

Black Forest Labs’ staged rollout separates current access from planned variants. The company says FLUX 3 Image will join the application program in the following weeks, with API and private-weight availability in later phases. An open-weight FLUX 3 Dev edition is also planned for later in 2026, giving developers downloadable parameters they can run or inspect.

FLUX 3’s phased expansion follows the image-focused FLUX.2 generation.

Robotics, Rivals and the Next Proof Points

FLUX-mimic carries the shared-backbone approach into robotics. Mimic Robotics worked with Black Forest Labs on the system, which translates intermediate video features into robot commands through a lightweight action decoder. Company tests used robot data to adapt FLUX-mimic to tasks, but outside researchers have not independently reproduced the result.

Video prediction dominates the reported training load. In company measurements, it consumed more than 95% of training compute, while audio represented less than 0.5% of tokens in a 720p clip with sound.

Audi’s tests move the robotics work from architecture to a factory setting. Company measurements put the FLUX-mimic backbone below 80 milliseconds on one Nvidia RTX 5090 and the complete system at 101 milliseconds. Audi tested and deployed FLUX-mimic on production and logistics work involving flexible seals and cables, electronic control-unit insertion, and parts kitting.

The release of the planned open-weight FLUX 3 Dev variant later in 2026 will determine when independent inspection and deployment can begin.

Markus Kasanmascheff
Markus Kasanmascheff
Markus has been covering the tech industry for more than 15 years. He is holding a Master´s degree in International Economics and is the founder and managing editor of Winbuzzer.com.
Subscribe
Notify of
guest
0 Comments
Newest
Oldest Most Voted
Inline Feedbacks
View all comments