Comparisons

    Animated stills or generated video for short ads

    A paced slideshow of generated stills beats generated video on revision cost and variants. It loses on anything that must be demonstrated. One rule decides.

    Versely Team9 min read

    The decision usually gets made on vibes: video feels more premium, so video wins. Then the brief comes back with three changes and you find out what a revision costs when your smallest editable unit is a five-second generation instead of a single frame. The two formats have genuinely different cost structures, and the difference isn't quality — it's what you have to regenerate when something is wrong.

    One question resolves most briefs: does the ad need to demonstrate motion, or does it need to deliver information? If a viewer has to see something move to believe the claim — a pour, a fold, a before-and-after, a hand applying the product — that's video and nothing else will do. If the ad is a list, a comparison, a claim stack, a testimonial quote or a step sequence, the motion is doing no work and you're paying for it anyway.

    The revision unit is the whole argument

    Every other difference follows from this one.

    A slideshow's smallest editable unit is one image. If slide four is wrong, you regenerate slide four. On the catalog that's Nano Banana 2 at a flat 4 credits per call, or Flux 2 Flash at 1 credit per megapixel. The other slides are untouched, the pacing is untouched, the overlays are untouched.

    A generated video's smallest editable unit is a clip. If the second shot is wrong, you regenerate the second shot — and on Vidu Q3 Image to Video that's 4 credits per second at SD with an 18-credit minimum per generation, meaning a two-second fix and a four-second fix bill identically. There's no partial regeneration of a clip; a bad half-second is a whole new generation.

    That gap is roughly an order of magnitude on a single fix, and it multiplies by the number of fixes a real brief produces. Three rounds of notes on a six-slide deck is a handful of image regenerations. Three rounds of notes on a six-shot video is a re-render of most of the ad.

    The tooling reflects this. The slideshow surface exposes slide-level operations directly — replace a single slide, remove one, reorder them, append more — so a note like "move the price reveal to the end and swap the opener" is a reorder plus one regeneration rather than a rebuild. There's no equivalent for a generated clip: you can trim it, speed it up or cut around it in the editor, but you cannot regenerate its middle.

    What the two formats actually cost

    The 19 published slideshows are all 9:16 reels, and their totals are on the record: six-image reels at 36 credits, seven-image reels at 42, and an eleven-image reel at 242. One six-image reel in that set costs 132 rather than 36 — same slide count, different image model. That spread is the real lesson: on the stills side, the model you pick is a bigger cost lever than the number of slides.

    On the video side, the published workflow recipes run from 60 credits for the shortest reel up to 1,800 for the largest multi-scene build at preview resolution. The overlap is real — a cheap video recipe and an expensive slideshow are in the same range — so "stills are cheaper" is not reliably true on the first pass. It becomes true on the third revision.

    Paced generated stills Generated video
    Smallest editable unit One image One clip
    Cost of a single fix One image generation One full clip generation, at the model's minimum
    Reordering the structure An edit, no regeneration A re-cut, no regeneration
    New hook, same body One image One clip
    Time to first playable cut Minutes Generation queue per scene
    Can show a transformation No Yes
    Motion is The cut cadence and any pan or zoom pass Inside the frame, generated
    Frame rate Set at export 25 fps by default

    Scroll-stopping: what each format can actually do

    Be careful with the received wisdom here, because "video stops the scroll" collapses two different mechanisms.

    The first is frame-one contrast — whether the opening frame is visually distinct from the feed above and below it. Both formats compete on equal terms for this, because both open on a still frame. A generated still gives you more control over it, not less: you compose exactly the frame you want, and you can regenerate it until it's right for the cost of one image.

    The second is the first motion beat — the moment something changes and confirms this isn't a static post. A slideshow's first motion beat is the first cut, and you set its timing to the frame. A generated video's first motion beat is whatever the model produced, which may be a slow drift for the first second. That's not a knock on video; it's a reason to check the opening second of a generated clip specifically, and to consider whether the cut you'd control is worth more than the motion you didn't.

    What a slideshow genuinely cannot do is demonstrate. A cut between "before" and "after" asserts a transformation. A continuous shot of the transformation shows it. For any claim where the proof is the movement itself, that difference is the entire ad and there is no pacing trick that substitutes.

    Two things carry across both formats and are worth getting right regardless: burned-in captions, because most of the feed is watched muted, and a measurable hook rate so the comparison is settled by data instead of preference.

    Variants per brief

    This is where the formats diverge most sharply, and it's the axis that decides the format for most testing work.

    From one slideshow you get a large number of legitimate variants without generating anything new: reorder the slides, cut it from eight down to six, change the overlay copy, swap the caption style, change the music and the transition timing. Each of those is an edit. Only a genuine content change — a new hook image, a different product shot — costs a generation, and it costs one image.

    From one video, the same list is much shorter. Reordering shots and re-cutting length are edits. Changing the hook is a new clip. Changing what happens inside a shot is a new clip. Every content-level variant is a generation at the model's minimum.

    There's a real lever on the video side that closes part of the gap: the editor's preview pass renders at 480p at no credit cost, subject to a short per-user cooldown, and charges once for the final export regardless of how many clips are on the timeline. So cutting fifteen versions of an assembly and exporting the one that wins is one export charge, not fifteen. The previews-and-export page works through the formula, and the editor is where that loop lives. It makes assembly variants cheap on both sides — it does nothing for variants that need new footage.

    The rule, and the hybrid

    Use paced stills when the ad's job is to deliver information: listicles, comparisons, claim stacks, testimonial quotes, step-by-step sequences, price and offer reveals, anything where the viewer is reading. The slideshow maker is the surface, the published set shows what the format looks like at 6 to 11 slides, and brand carousels covers pushing it toward branded work.

    Use generated video when the claim is physical: application, texture, pour, fold, fit, assembly, a before-and-after that has to be continuous, or a person delivering a line where the delivery is the point.

    The hybrid that usually wins is neither: a paced stills structure with one generated video shot dropped in where the demonstration happens. You get slide-level revision economics for the 80% of the ad that's information, and one real motion shot for the moment that has to be shown. Structurally it's a slideshow converted to video with one clip spliced on the timeline — transitions, music and export covers the conversion step, and the assembly happens in the same editor either way.

    Start from stills unless the brief contains a demonstration. That's the default, and the exceptions announce themselves.

    FAQ

    Is a slideshow ad obviously a slideshow to a viewer?

    Only when the pacing gives it away. Slides held at a uniform interval read as a template; varying the hold time against the copy length and the audio does not. The tell is rhythm, not resolution, which is why the timing pass is worth more attention than the image quality pass on this format.

    Should the caption sit on the image or over it?

    Both approaches exist and they're different jobs. Text rendered into the image by the image model is baked into the composition and survives any re-crop, but changing the copy means regenerating the slide. An overlay added afterwards is editable without touching the image underneath it. For copy that will change across variants, overlay. For copy that is part of the design, render it in.

    How many slides is right for a short ad?

    The published set clusters at six to seven, with one at eleven, and that band is a reasonable starting point for a 9:16 reel. The real constraint is reading time: every slide has to be legible at the hold length you've set, so the slide count follows from how much copy each slide carries rather than from a target duration.

    Can I test both formats against the same brief without doubling the budget?

    Yes, and it's the honest way to settle it. Build the stills version first, because it's the cheaper structure to revise, and get the copy and ordering right there. Then generate only the one shot that would justify the video version and splice it in. You end up with two testable assets that share a script, and the difference between their results tells you whether the motion was doing any work.