AI News

    FLUX 3 makes Black Forest Labs a video lab

    Black Forest Labs trained FLUX 3 jointly on image, video and audio with 20s synced sound. Which Flux still-image workflows should absorb it, and which won't.

    Versely Team7 min read

    For two years Black Forest Labs had one job in most people's mental model: the Flux image line. Fast stills, strong prompt adherence, a credible open-weight story. If you needed motion you went somewhere else and came back for the frames.

    That framing stopped being accurate on 23 July 2026, when BFL published FLUX 3 on its own blog. Not a video add-on bolted to an image backbone — one model trained jointly across image, video and audio, generating video with native synchronised sound up to 20 seconds. API access and private weight access shipped at launch, with an open-weight "FLUX 3 Dev" backbone promised later.

    A creator workspace with multiple screens, cameras, and editing tools

    What "jointly trained" actually implies

    The phrase does real work here. A model that generates sound as a separate pass is doing dubbing, and dubbing has a sync problem you can hear. A model that learned picture and audio together produces a soundtrack that agrees with what's on screen because it was never a separate object. That's the difference between a clip with music laid over it and a clip where the footstep lands on the footstep.

    Twenty seconds is the other number that matters. Most of the field was quoting 5 to 10 seconds a year ago, and the honest workaround was stitching. Twenty seconds of continuous native audio is long enough for a full product beat, a complete piece of VO, or a single unbroken demo without a cut hiding a seam.

    The strategic read: BFL is now competing on the same board as the video labs. That means when you evaluate the Flux family, "which Flux image model" is no longer the whole question.

    Where FLUX 3 sits in the Versely catalog

    This is the part that surprises people who know Flux only as an image line. Inside Versely, FLUX 3 is a video family, not an image one:

    Model What it does Duration Credits for a clip
    Flux 3 Text to Video Video with native audio from a prompt alone, eight aspect ratios 5–20s 43–170
    Flux 3 First Last Frame to Video Interpolates between a defined start and end frame, with audio 5–20s 43–170
    Flux 3 Extend Video Continues an existing clip past its final frame 5–20s 103–410

    Read those ranges carefully, because the video side of the family bills by the second rather than by the generation. The low end of each range is a 5-second clip and the high end is the full 20. That is a different mental model from the stills, where a generation is a generation and the number does not move with a duration slider. Budget video by seconds of finished footage, not by number of attempts.

    The Flux stills you have been using are a different generation of the family. Flux 2 Max at 7 credits, Flux 2 Flex at 5, Flux Kontext at 4 for instruction-based edits, Flux 1.1 Pro at 4, and the 1-credit fast tiers. Those are flat per-image prices and they are unaffected by FLUX 3 landing. Nothing about your still pipeline breaks. What changes is that the motion step no longer has to leave the family.

    Four Flux still workflows that should absorb it

    Not every image workflow wants video attached. These four do, and the reason in each case is that you were already faking motion or already leaving Flux to get it.

    1. The hero still that gets a "can we get a moving version" request. This is the most common one. You generate a hero frame in Flux 2 Max, the client asks for a five-second loop for the site header, and you re-prompt in a different model and spend an hour matching the look. Instead: take the approved frame, pin it as the start, pin a slightly different composition as the end, and run first/last frame. The grade carries because you supplied both endpoints.

    2. Product stills that need a turn, not a scene. A packshot rotating through 20 degrees is a shot, not a film. First/last frame handles it and the audio pass gives you a soft room tone instead of dead silence, which matters more than people expect on autoplay-muted platforms where a user un-mutes.

    3. Social statics that keep getting cut down into Reels. If your team is already exporting stills and animating them in an editor with a Ken Burns push, that's a workflow FLUX 3 absorbs outright. One generation replaces still-plus-manual-motion, and it comes with sound.

    4. Concept boards that need a motion test before sign-off. Boarding in stills is cheap and fast, but a client approving a static frame is approving something they can't quite picture moving. A 5-second text-to-video pass on the approved board direction closes that gap before you commit to the expensive version.

    What should stay still

    Three categories where pulling video into the workflow is a downgrade:

    • Anything with legible on-screen type. Video models garble text worse than image models do, and 20 seconds gives the garbling 20 seconds to wobble. Generate the plate, add type in post.
    • Print and large-format. A video frame is not a print-resolution still, whatever the pixel count says. Keep the stills pipeline for anything going to press or to a large physical surface.
    • Volume catalogue work. Two hundred product variants is an image job. It stays an image job. The 1-credit Flux fast tiers exist precisely for that shape of work, and putting motion on all two hundred is a budget decision nobody asked for.

    How to actually run the switch

    A concrete sequence rather than a vibe:

    1. Pick one live workflow where a still currently gets animated by hand.
    2. Regenerate the approved frame set from your existing Flux image prompt, unchanged.
    3. Run the same beat three ways: text-to-video cold, first/last frame from your approved endpoints, and your existing manual animation. Same brief, same aspect ratio.
    4. Judge them muted first, then with sound. The audio pass is the thing you cannot get from the manual route, so evaluate it separately or it will bias the picture comparison.
    5. If first/last frame wins, and in endpoint-controlled work it usually does, make it the default and keep text-to-video for the cases where no approved frame exists yet.

    The Flux 3 prompting guide covers the soundscape and staging language that actually moves the audio output, which is the part most people under-specify on the first attempt. If you want the family view rather than the model view, Black Forest Labs' provider page lists everything of theirs currently carried.

    FAQ

    Is FLUX 3 an image model or a video model?

    Both, by design. BFL trained one model jointly on image, video and audio rather than shipping a video head on an image backbone. In practical terms, what Versely carries under the FLUX 3 name is the video side: text-to-video, first/last frame, image-to-video and extend, all with native synchronised audio.

    Does FLUX 3 replace the Flux image models I already use?

    No. Flux 2 Max, Flux 2 Flex, Flux Kontext and the fast tiers are unchanged and still the right tools for stills, edits and volume catalogue work. FLUX 3 changes what happens at the motion step, not the still step.

    How long can a FLUX 3 clip run?

    Up to 20 seconds with synchronised audio, selectable in one-second increments from 5s upward. Flux 3 Extend Video continues an existing clip past its final frame if you need to go further, though the source clip has its own limits on size and length.

    Is there an open-weight version of FLUX 3?

    BFL announced an open-weight "FLUX 3 Dev" backbone as a later release. At launch, access was API plus private weight access. Announced is not shipped, so plan around the hosted path and treat the open backbone as an option that may arrive rather than a date you can build a schedule on.

    Run the comparison yourself in the AI video generator. The same brief through Flux 3 Text to Video and your current motion model takes about ten minutes, and it settles the argument better than a launch post will.