Flux 3: everything about the new multimodal model

Flux 3: everything about the new multimodal model

Black Forest Labs has officially unveiled Flux 3. The company, which became well known for its strong image models FLUX.1 and FLUX.2, is now taking on video, audio, and even action prediction for robotics.

FLUX 3 tries to unify things that used to live separately: images, motion, sound, references, text on images, and physical causality. Right now it's early access, not a full public launch. Let's break down what to expect from the new model.

What is Flux 3

Flux 3 is Black Forest Labs' new multimodal model for visual content. In the official release, BFL writes that the model is jointly trained on images, video, and audio, because no single modality on its own fully describes the world.

The idea is simple but technically strong:

  • Images capture the structure of a scene at a single moment in time.
  • Video adds dynamics, motion, object contact, and cause-and-effect relationships.
  • Audio helps understand what's happening in the frame: an impact, a footstep, speech, mechanical noise, the atmosphere of a space.
  • Language ties all of this to the user's task.

Instead of a set of separate models, Black Forest Labs is betting on a single backbone that should learn a more coherent representation of the world. That's why FLUX 3 is positioned not just as a creative model, but as a foundation for physical AI.

Self-Flow: the technical foundation of Flux 3

FLUX 3 is built on the Self-Flow approach. It's an extension of flow matching, where generative training is combined with learning useful internal representations.

In standard diffusion/flow models, the task often boils down to local reconstruction of noisy data. That works well for generation, but doesn't always force the model to build strong global semantics. That's why many approaches used external encoder models, such as DINO, to "feed" the generator richer features.

Self-Flow tries to remove this dependency. Its key mechanism is Dual-Timestep Scheduling: different tokens receive different levels of noise. Part of the input stays relatively cleaner, part is more heavily corrupted, and the model is forced to reconstruct the missing information from context. This creates information asymmetry that pushes the model to learn not just pixels, but scene structure.

For FLUX 3 this matters because the same architecture has to work across different modalities. If the internal representations improve, that helps not only image or video generation, but downstream tasks like action prediction too.

What Flux 3 Video can do

The loudest part of the release is FLUX 3 Video. According to Black Forest Labs, the model can generate video with native audio up to 20 seconds long in a single pass.

Claimed capabilities:

  • text-to-video: generating a clip from a text prompt;
  • image-to-video: animating a starting frame or using an image as a visual reference;
  • video-to-video: carrying over key elements from a source clip, such as a character, into a new context;
  • video/audio continuation: continuing input video and audio;
  • keyframe-to-video: generating a transition between given keyframes;
  • multilingual dialogue: multilingual speech;
  • different styles and aspect ratios, not just a "cinematic" look;
  • agentic chaining: linking individual clips into longer multi-shot sequences;
  • typography generation and animated designs.

The most important part here is that sound isn't added as separate post-processing. BFL talks about native audio generation: the model generates video and audio together, so the sound should match physical events in the frame more closely.

What about images

Although the release is discussed loudest as video news, FLUX 3 remains important for images too. Black Forest Labs states that the model will be able to synthesize and edit images in different styles, aspect ratios, and resolutions.

According to BFL's early midtraining evaluations, FLUX 3 handles complex prompts and text generation better than previous versions. More accurate multilingual text rendering is mentioned separately. That matters for ads, posters, interfaces, infographics, and any task where an image needs to be not just pretty, but functional.

However, FLUX 3 Image isn't broadly open yet. Black Forest Labs promises to launch early access for image use cases in the coming weeks.

Action prediction and FLUX-mimic

The most unusual part of the release is FLUX 3's connection to robotics. Black Forest Labs and mimic robotics have unveiled FLUX-mimic, a video-action model built on top of FLUX 3.

The logic goes like this: to predict video well, a model has to understand how objects move, collide, deform, and affect each other. That's not full physics in the academic sense, but it's a useful approximation of the world. If the model has already learned that dynamics, its internal representations can be used to predict robot actions.

The BFL x mimic technical paper has some interesting details. The company writes that video prediction accounts for more than 95% of compute costs when training FLUX 3, while audio in 720p video with audio makes up less than 0.5% of tokens. In other words, in their view of the world, video is the heaviest modality, and audio and robot actions become additional projections of a single physical reality.

mimic robotics claims that FLUX-mimic is already being tested with Audi on real production tasks: kitting, inserting electronic components, assembly, and working with soft, flexible parts like gaskets and cables. In a preview soft-body kitting benchmark, mimic claims a 95% success rate versus 55% for an adapted pi0.5 and 70% for a flow matching baseline policy.

That sounds strong, but for now it's partner-provided data. A final verdict needs independent verification, a description of test conditions, baseline systems, and statistics from real production lines.

Early quality assessments

Black Forest Labs published preliminary preference scores for video. In its tests, the company generated 10-second clips in 720p with audio and compared FLUX 3 against other video models.

According to BFL, FLUX 3 was preferred over:

  • Grok Imagine Video — up to 69%;
  • Kling v3 Pro — 60%;
  • Happy Horse v1 — 59%;
  • Happy Horse 1.1 — 57%;
  • Seedance 2.0 and Gemini Omni Flash — 52%;
  • Runway Gen-4.5 — 77%;
  • Luma Ray 3.2 — 93%.

These numbers should be read carefully. A preference test shows which output raters preferred more often in a pairwise comparison, but it's not a universal benchmark. Until BFL discloses the full protocol, the prompt set, the number of comparisons, the cost per generation, and failure cases, these percentages are better treated as an early signal than final proof of leadership.

What's already available

The release is rolling out in stages:

  • FLUX 3 Video: early access by application, video + optional native audio;
  • FLUX 3 Action / FLUX-mimic: access through selected research and commercial partners;
  • FLUX 3 Image: early access promised in the coming weeks;
  • FLUX 3 Dev: an open-weight multimodal backbone promised later in 2026.

There's no public pricing for FLUX 3 at the time of writing. Black Forest Labs' pricing page currently shows FLUX.2 and separate Flux Tools, not distinct SKUs for FLUX 3 Video or FLUX 3 Image. BFL's documentation is also still mainly focused on FLUX.2, FLUX.1 Kontext, and editing tools.

Main advantages of Flux 3

The first advantage is a unified multimodal architecture. If video, image, and sound are trained together, the model potentially links motion, appearance, and acoustics better. For users, this can mean less manual pipeline assembly: no need to generate an image separately, animate it separately, synthesize sound separately, and then chase synchronization.

Second is native audio. For AI video, sound is often a weak point: it's added after generation and doesn't always match the action. FLUX 3 bets on joint generation, including speech, ambient sound, and effects tied to physical events.

Third is working with references and multi-shot logic. Image-to-video, video-to-video, keyframes, and clip chaining matter for real production work. Teams rarely need one random pretty clip; more often they need a character, product, style, or scene that stays consistent across variations.

Fourth is typography and design. Black Forest Labs specifically highlights typography generation and animated designs. If this holds up in practice, FLUX 3 could be useful not only for "cinematic shots" but also for ad creatives, intros, presentation videos, and interface graphics.

Fifth is a planned open-weight Dev version. That's critical for the FLUX community. Previous FLUX Dev models became important precisely because they could be run, adapted, and embedded locally. But for FLUX 3 Dev, the license, size, hardware requirements, and quality level relative to the closed version are still unknown.

What people who've already tried it are saying

Third-party reaction is mixed so far: a lot of interest, but also a lot of caution.

Decrypt calls FLUX 3 BFL's first video-model step and focuses on the robotics tie-in. At the same time, the outlet fairly points out that BFL's published percentages are a preference test, not a rigorous independent metric.

Magica offers a more skeptical take. Its main point: FLUX 3 is still gated early access, and the robotics angle looks promising but needs proof. For video, reproducible evaluations are needed: the exact model version, test set, rater protocol, resolution, cost of generation, and reliability of multi-shot workflows.

Novoads makes a useful point for ad teams: FLUX 3 is announced, not launched. In other words, the model has been unveiled, but it hasn't yet become an everyday tool you can plug into a production process with clear pricing and an SLA.

Aireiter writes that early testers on launch day posted clips and cited practical parameters like 20 seconds, up to 10 reference media, and aspect ratios up to 21:9. But BFL hasn't locked these details in as formal specs, so they should be treated as user reports rather than a guarantee.

On r/StableDiffusion, the reaction is livelier and more practical. Users praise individual examples, are waiting for open weights, and are asking the key questions: how much VRAM/RAM will be needed, will the Dev version be full-featured, how to verify the provenance of published clips, and how much FLUX 3 actually beats Seedance 2.0 in complex scenes.

AI newsletters like The Rundown and Latent Space greeted the release noticeably more optimistically. For them, the main story isn't any single clip, but the fact that Black Forest Labs is finally expanding FLUX into audio-video and physical AI, and a future open-weight Dev version could make this architecture accessible to researchers and developers.

What this means for the market

FLUX 3 arrives as AI video is maturing fast. It's no longer enough to just make one pretty short clip. Users need references, character consistency, precise actions, sound, dialogue, different formats, controllability, repeatability, and clear pricing.

If FLUX 3 lives up to its claimed capabilities, it could become one of the main competitors to Seedance, Kling, Runway, Luma, Veo, and Sora-like models. Especially in cases where what's needed isn't a single "wow clip" but a pipeline for regular generation: advertising, e-commerce, social video, product motion, previz, motion design, and scene prototyping.

For GPTunneL users, the main practical takeaway is this: FLUX 3 is worth watching not just as a new video generator, but as a possible new class of multimodal models where image, sound, and motion all emerge from a single architecture.

Bottom line

Flux 3 is a major release for Black Forest Labs. The company clearly wants to move from static images to a broader visual intelligence stack: video with native sound, images, editing, long sequences, physical dynamics, and action prediction.

The most interesting parts are Self-Flow, native audio, multi-shot workflows, the claimed strong typography, and the FLUX-mimic tie-in. The weakest points right now are the early-access stage, lack of pricing, incomplete benchmark methodology, and unknown requirements for the open-weight Dev version.

So it's better to view FLUX 3 not as a finished product that's already proven itself, but as a strong pitch. If Black Forest Labs opens up access, publishes its methodology, and ships a genuinely usable Dev version, Flux 3 could become one of the key releases of 2026 in video, image, and multimodal AI generation.