MarkTechPost→ original

Black Forest Labs выпустила FLUX 3: видео до 20 секунд со звуком из одной модели

Black Forest Labs выпустила FLUX 3 — мультимодальную flow-модель, обученную сразу на изображениях, видео и звуке. Из одного набора весов она генерирует картинки, видео до 20 секунд с нативным аудио и предсказывает действия робота. В тесте на человеческих предпочтениях FLUX 3 обошла Luma Ray 3.2 в 93% сравнений и Runway Gen-4.5 в 77%. Тот же backbone управляет роботом быстрее 80 мс на одной RTX 5090.

AI-processed from MarkTechPost; edited by Hamidun News
Black Forest Labs выпустила FLUX 3: видео до 20 секунд со звуком из одной модели
Source: MarkTechPost. Collage: Hamidun News.
◐ Listen to article

Black Forest Labs released FLUX 3 on July 26, 2026 — a multimodal flow model that generates images, video up to 20 seconds long with native sound, audio, and predicts robot actions, all from a single set of weights. This is the first model in the FLUX family where video, sound, and action prediction emerge from a shared architecture.

What FLUX 3 Video can do

FLUX 3 Video creates clips up to 20 seconds long in a single generation, with sound included from the start. According to Black Forest Labs, the model supports five modes: text-to-video, image-to-video, video-to-video with a reference, keyframe-to-video for controlled transitions, and continuation of video-audio from an input clip.

The company also highlights multilingual dialogue, the assembly of several clips into a single multi-scene sequence, and the generation of animated typography. The BFL team says it has strongly refined human facial expressions and precise syncing of sound to physical events — a hit landing on screen coincides with the sound of the impact.

  • Release — July 26, 2026
  • Video up to 20 seconds with native sound in a single generation
  • Five modes: text/image/video/keyframe-to-video and video-audio continuation
  • FLUX-mimic: the robot policy runs faster than 80 ms on a single RTX 5090
  • Rollout is staged: Video and Action get early access, Image comes later, open weights last

How the model is built

FLUX 3 is built on the Self-Flow method — a Black Forest Labs approach that combines flow matching with self-supervised feature reconstruction in a single architecture. The company introduced Self-Flow itself back in March 2026; what's new here is the scale: according to BFL, compute and data volume were "significantly increased," training the model on video, images, and sound simultaneously.

Resource allocation is uneven: according to Black Forest Labs, video prediction consumes over 95% of training compute, while audio accounts for less than 0.5% of tokens. The idea is that the modalities constrain each other during training.

"No single modality gives a complete description of the world," is how

Black Forest Labs' research team formulates its principle.

How much does FLUX 3 outperform competitors

In a preliminary human-preference test, FLUX 3 beat most of its rivals. The comparison used 10-second text-to-video clips at 720p with sound.

  • Luma Ray 3.2 — FLUX 3 was preferred in 93% of comparisons
  • Runway Gen-4.5 — in 77%
  • Grok Imagine Video — up to 69%, Kling v3 Pro — 60%
  • Seedance 2.0 and Gemini Omni Flash — 52%, nearly a coin flip

The same backbone powers FLUX-mimic — a robot policy that, according to BFL, runs faster than 80 ms on a single RTX 5090 GPU. This shows that the multimodal model applies not only to media generation but also to controlling physical actions.

What this means

FLUX 3 is a bid for a unified foundation model, where image generation, video with sound, and robot control all grow out of shared weights. Access is still limited: Video and Action are open in early access, Image will arrive later, and Black Forest Labs says it will release the open weights last.

Frequently asked questions

What is FLUX 3?

FLUX 3 is Black Forest Labs' multimodal flow model, trained simultaneously on images, video, and sound. From a single set of weights it generates images, video up to 20 seconds long with native audio, and predicts robot actions.

When will the open weights for FLUX 3 be released?

There's no exact date yet. Access is rolling out in stages: first Video and Action in early access, then Image, while Black Forest Labs plans to publish the open weights last.

What powers the robot built on FLUX 3?

FLUX-mimic — a robot policy built on the same backbone. According to Black Forest Labs, it runs faster than 80 ms on a single NVIDIA RTX 5090 GPU.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…