Skip to content

FLUX 3 folds image, video, audio, and robot actions into one model

Black Forest Labs released FLUX 3 on July 23, and it's a bigger jump than the version number suggests. The Freiburg lab built its name on FLUX, the image models that generate art inside Photoshop, Canva, Picsart, and, closer to home, inside Cinevva for everyone who picks the Flux provider. FLUX 3 is the first time the company ships video at all, and it doesn't ship a separate video model. It trains one set of weights on images, video, and audio at once, then extends that same backbone to predict robot actions. One model, four kinds of output.

A first look at what FLUX 3 Video generates

One backbone, not four products glued together

The pitch rests on a research idea BFL calls Self-Flow: instead of learning image generation, video generation, and audio generation as separate problems, learn a single representation of how the world works. An image is a slice of space at one moment. Video adds time and motion. Audio carries the causal link between an impact and the sound it makes. Train on all of them together and the constraints reinforce each other, so the sound has to match the motion and the motion has to obey physics. The surprise in the launch is that action prediction fell out of the same backbone without wrecking the video quality, which is why FLUX 3 can drive a robot arm and generate a music video from the same weights.

The video, concretely

FLUX 3 Video makes clips up to 20 seconds long in a single pass, with native audio generated alongside the picture, dialogue, sound effects, and ambient noise included. That length matches OpenAI's now-shut-down Sora and beats most of the current field on duration. It handles text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue, and agentic chaining of clips into multi-shot sequences. In BFL's own early tests on 10-second 720p clips, reviewers preferred it over Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77%. Against the strongest rivals the margin shrinks to a coin flip: 52% over both ByteDance's Seedance 2.0 and Google's Gemini Omni Flash. These are the lab's own preliminary numbers, so treat them as a claim, not a verdict.

The catch is access. FLUX 3 Video is in gated early access through the API and private weights, and you have to apply and get approved. FLUX 3 Action is partner-only, already being tested on a real Audi production line through robotics startup mimic. FLUX 3 Image, the tier most designers actually want, lands in the coming weeks. The open-weight FLUX 3 Dev backbone, the piece the local-inference crowd cares about, isn't due until later in 2026.

Why an image lab going multimodal matters to us

We watch BFL closely because Cinevva already runs on their work. When one of the two labs whose image models sit inside the world's biggest creative tools decides its future is a unified video-audio-action model, that's a signal about where the whole creation stack is heading, not just a new toy. The line between "generate a picture," "generate a clip," and "simulate a world you can act in" is dissolving into one model, and it's the same convergence we flagged with the MIRA world model from the other direction.

For creators the practical read is steady. A better, longer, audio-native video model lowers the cost of the first draft again, same as Seedance, Kling, and Gemini did before it. What it doesn't touch is the part that was always hard: knowing what's worth making, then iterating a generated clip into something people actually want to watch or play. That's the loop we keep building around, from the prompt to sharing, remixing, and publishing on the open web. FLUX 3 makes the raw material cheaper and better. Turning that material into a game or a story people care about is still the work.

References