FLUX 3 Video
Black Forest Labs (Freiburg, Germany) — one multimodal model where picture and sound come out of the same request. 5 to 20 seconds without a cut, multilingual dialogue with lip-sync, and a storyboard of up to ten images the camera travels through.
Strengths
20 seconds without a single cut
Other models give you five to fifteen seconds, then you cut. FLUX 3 renders up to twenty seconds in one go — with a continuous audio track, no cut seams, no continuity errors.
German dialogue, word for word and lip-synced
Audio is produced in the same request as the picture, multilingual and lip-synced. We checked it with Whisper: the German line we specified came back word for word, with none of the mumbled pseudo-German other models produce.
Storyboard: the camera travels, it does not cut
Give it up to ten images with a second each, and FLUX travels from image to image in one take. No morphing, no hard cut — measured on real client material: two product photos became one continuous move that lands exactly on the target frame.
Cinema ratios at no extra cost
21:9 and 2:1 are beyond every other video model we run. With FLUX the price follows pixels per frame, not the aspect ratio — so Cinemascope costs exactly what 16:9 costs.
Showcase
Real generations
Best for
Testimonials and presenter segments in one take
Product spots travelling from one approved frame to the next
Long social clips without visible cut seams
Cinemascope shots for brands working in cinema formats
Bringing an already approved still into motion
Not its strong suit
No subject references: FLUX 3 cannot do "put this character or product into the scene". It only knows temporal keyframes. For subjects in free scenes use Kling 3.0 Omni, Seedance or Gemini Omni — Black Forest Labs has announced references for a later release.
Small on-image text is re-rendered and letters shift in the process. Either avoid frames with fine print, or add type in post.
Full HD maximum (1920×1088), no 4K — and nothing shorter than five seconds.
Technical details
Architecture
FLUX 3 comes from Black Forest Labs in Freiburg — the team behind the latent diffusion architecture used in Stable Diffusion. Instead of separate image and audio models, a single set of weights was trained on image, video and audio at the same time. That is why the sound is not added afterwards but emerges together with the frames. Only the video capability is exposed through the API so far; the image model has been announced but not shipped.
Key features
- 5 to 20 seconds in one take, 24 frames per second
- Native audio in the same request — multilingual, lip-synced
- Storyboard of up to 10 keyframes, optionally with exact seconds
- Start and end frame as the simple case of a storyboard
- Seven aspect ratios: 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16
- HD and Full HD — priced by pixels per frame, not by aspect ratio
Ready to get started with FLUX 3 Video?
Beta access is free for selected creators. Claim your slot.
More models