freycreate
By Black Forest Labs·23 July 2026

FLUX 3 Video

Black Forest Labs (Freiburg, Germany) — one multimodal model where picture and sound come out of the same request. 5 to 20 seconds without a cut, multilingual dialogue with lip-sync, and a storyboard of up to ten images the camera travels through.

Resolutionup to 1080p
Duration5-20 sec
Audionative, lip-synced
Keyframesup to 10
Aspects7 ratios
ReleaseJuly 2026

Strengths

20 seconds without a single cut

Other models give you five to fifteen seconds, then you cut. FLUX 3 renders up to twenty seconds in one go — with a continuous audio track, no cut seams, no continuity errors.

German dialogue, word for word and lip-synced

Audio is produced in the same request as the picture, multilingual and lip-synced. We checked it with Whisper: the German line we specified came back word for word, with none of the mumbled pseudo-German other models produce.

Storyboard: the camera travels, it does not cut

Give it up to ten images with a second each, and FLUX travels from image to image in one take. No morphing, no hard cut — measured on real client material: two product photos became one continuous move that lands exactly on the target frame.

Cinema ratios at no extra cost

21:9 and 2:1 are beyond every other video model we run. With FLUX the price follows pixels per frame, not the aspect ratio — so Cinemascope costs exactly what 16:9 costs.

Showcase

Real generations

Best for

Testimonials and presenter segments in one take

Product spots travelling from one approved frame to the next

Long social clips without visible cut seams

Cinemascope shots for brands working in cinema formats

Bringing an already approved still into motion

Not its strong suit

No subject references: FLUX 3 cannot do "put this character or product into the scene". It only knows temporal keyframes. For subjects in free scenes use Kling 3.0 Omni, Seedance or Gemini Omni — Black Forest Labs has announced references for a later release.

Small on-image text is re-rendered and letters shift in the process. Either avoid frames with fine print, or add type in post.

Full HD maximum (1920×1088), no 4K — and nothing shorter than five seconds.

Technical details

Architecture

FLUX 3 comes from Black Forest Labs in Freiburg — the team behind the latent diffusion architecture used in Stable Diffusion. Instead of separate image and audio models, a single set of weights was trained on image, video and audio at the same time. That is why the sound is not added afterwards but emerges together with the frames. Only the video capability is exposed through the API so far; the image model has been announced but not shipped.

Key features

  • 5 to 20 seconds in one take, 24 frames per second
  • Native audio in the same request — multilingual, lip-synced
  • Storyboard of up to 10 keyframes, optionally with exact seconds
  • Start and end frame as the simple case of a storyboard
  • Seven aspect ratios: 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16
  • HD and Full HD — priced by pixels per frame, not by aspect ratio
Use directly from AI agents (MCP)

Ready to get started with FLUX 3 Video?

Beta access is free for selected creators. Claim your slot.

More models

Also interesting

We use cookies 🍪 ·