Guides

    Shooting a Launch Video Without a Camera

    No camera, no studio, no on-camera talent, and in some cases no product yet. A launch video plan that names the substitute for every missing input.

    Versely Team9 min read

    Most advice about making a launch video assumes you are missing budget. The harder version is missing inputs: nothing to film with, nowhere to film, nobody to put on screen, and — more often than anyone admits — no physical product yet, because the launch date landed before the samples did.

    That is solvable, but only if you treat each missing input as a separate substitution rather than one big absence. Of the 296 models in Versely's catalog, 168 require neither an uploaded image nor uploaded footage: a written prompt is the only input they need. 113 refuse to start without a still, 22 refuse to start without video, and 7 of those insist on both. Knowing which lane you are in is the difference between a plan and a week of frustration.

    Name the five inputs you are missing

    A conventional launch shoot needs five things. Take them one at a time, because the substitute is different for each and mixing them up is how people end up trying to solve a talent problem with a lighting fix.

    Missing input The substitute Where the routing lives
    A camera Text-only generation — the prompt is the footage Making videos without a camera
    A studio Either generate the whole environment, or shoot one still and cut the room out Making videos without a studio
    On-camera talent A generated presenter, a voice over B-roll, or a recurring cast you own Making videos without showing your face
    A voice actor A generated read, a cloned voice, or two generated voices in conversation Making videos without a voice actor
    The product itself Generate the pack shot first and lock it as a reusable asset Making videos without a product sample

    The order in that table is not arbitrary. Talent and product are asset decisions; camera and studio are generation decisions. Assets have to be locked before generation starts, because every clip afterwards inherits them. Get that backwards and you will generate eleven beautiful shots of eleven subtly different products.

    Day 1 — lock the assets, generate nothing

    The single biggest failure in no-camera production is continuity, and continuity does not come from repeating adjectives in every prompt. It comes from reusing named reference assets.

    So day one produces no video at all. It produces:

    1. The pack shot. If the product does not physically exist, generate it. If one supplier photograph exists, that is your seed and everything else derives from it. Either way you end up with one canonical image of the product, and every subsequent clip conditions on it.
    2. The presenter, if there is one. A generated face that appears in shot one has to be the same face in shot six. Lock it as a named asset now.
    3. The environment. If more than one clip is set in the same place, decide what that place looks like once. Building the set once and reusing it is cheaper and more consistent than describing it eleven times.

    Setting up reusable characters and products is the mechanism. This is also the step where a launch with no physical sample stops being a blocker: you need a product image before you need a product video, and once the pack shot is locked, the absence of a sample no longer constrains anything downstream.

    Write your one-sentence premise on day one too — who, what changes, how it ends. Scene splitting works from a premise. It does not work from a mood.

    Day 2 — the shot list, and the one prompt discipline that matters

    A launch video is usually six to twelve beats. Write each as one continuous shot: subject, what it does, camera move, lighting, look. One shot, not a sequence — text-to-video models cut badly when a single prompt asks for two locations, and a "why did it jump" result almost always traces back to a prompt that requested a change of place.

    Set the aspect ratio before you generate, never after. Vertical output framed for 16:9 loses the subject's head, and reframing afterwards is a crop, not a fix.

    For a product launch, the shot types that actually carry weight:

    • The hands-holding-product POV. The most reliable product shot in short-form, because it establishes scale and reality in one frame without needing a face. The POV hands-product-hold prompt format is the ready-made version.
    • The problem beat. Whatever the product removes, shown rather than described.
    • The reveal. Product enters frame with the pack shot as the reference input.
    • The one detail shot. Texture, mechanism, or the thing on the label that matters.
    • The outcome. For anything digital or unphotogenic, this is the whole video — sell the change the product causes, not the object.

    Take two or three passes on each prompt before you rewrite it. Variance between takes on one prompt is usually wider than the difference between two prompts, and rewriting a prompt that would have worked on the next attempt is the most common way to waste a morning. When a take has the right motion but wrong details, keep it: details are fixable downstream, motion is not.

    Day 3 — the voice, decided before anything is cut

    The voice track is produced separately and mixed over the finished cut, which means it can be built in parallel with the shot list — but the decision has to come first, because it determines whether you need a talking face at all.

    Three routes:

    • One read. Type the script, get the narration. This is the default for a launch video and it needs no face on screen anywhere.
    • Two voices. If the format is a conversation — founder and customer, interviewer and expert — you need distinct speakers, and each speaker alias in the script has to map to a voice.
    • The same read in another language. If the launch is multi-market, dub the finished video rather than recasting it. The voice-over surface covers the language side.

    The decision that saves the most work: skip lipsync unless a face is genuinely on screen. Narration over product footage, B-roll or generated environments needs none, and lipsync is the most expensive job in the audio cluster.

    If you want the read in your own voice without ever being on camera, clone it once from a clean sample and reuse the voice ID forever. Sample quality is what determines clone quality, so record it somewhere quiet rather than fixing it later.

    Day 4 — assemble, and use the free pass properly

    Stitch the clips, then lay the voice and captions over the finished cut. Do not bake narration into individual clips — the moment you re-time the picture, baked-in narration forces you to regenerate.

    Iterate the assembly with the editor's preview mode, which renders a free 480p pass subject to a short per-user cooldown. The charged export happens once, on the cut you confirm, regardless of clip count. That is what makes a four-day plan viable for one person: pacing, caption timing and where the product reveal lands are all free to argue about, and only the answer costs anything.

    Captions last, after the cut length is settled, because the cut is what decides where lines break.

    What this actually costs, in credits

    The honest reference points are the published workflow recipes, because each one is a real multi-scene video with a real total. A few in the launch-adjacent shape:

    Workflow Scenes Preview render Full-resolution render
    Hivewell Raw Honey 10 310 credits 620 credits
    Plush Lipstick Influencer Review 3 186 credits 372 credits
    AlignPro — Posture Corrector Ad 15 451 credits 900 credits

    Those totals are the useful anchor because they cover every scene, not one clip in isolation. Note the pattern: on multi-scene work the preview-resolution total is roughly half the full-resolution one, which is why the sensible sequence is to build and approve at preview resolution and spend full resolution once. Every generation costs credits — there is no zero-credit lane anywhere in this plan except the editor's 480p preview pass, with its cooldown.

    Cloning a published recipe and swapping the asset map is also faster than building from a blank page, and it is the honest recommendation for a first launch video: pick the one whose structure matches yours, replace the pack shot and the script, and you have skipped the two hardest days.

    FAQ

    Does the launch video need a person in it at all?

    Only if the content is direct address — a pitch, an explanation, advice delivered to camera. If the information is the point, nobody on screen is not a limitation; product walkthroughs, problem-outcome cuts and narrated demos work fine with a voice and no face, and they remove both the talent problem and the lipsync cost.

    What if the product genuinely does not exist yet?

    Generate the pack shot and lock it as a reusable asset. That image becomes the canonical product for every downstream clip, so the physical sample stops being a dependency for the video even though it is still a dependency for shipping. When the real product arrives, swap the reference asset and regenerate the shots that show it — the script, cut and voice track all survive.

    How many takes should I budget per shot?

    Two or three on the identical prompt before changing anything, then one more after a single deliberate change if all of those failed the same way. On a ten-shot launch video, plan for meaningfully more generations than shots. Budgeting one render per shot is the assumption that makes people conclude AI video does not work after one attempt.

    Can I mix real footage with generated shots?

    Yes, and it usually helps. If you have any usable footage — even phone footage of the product on a desk — it can anchor the cut while generated shots carry the rest. The catch is consistency of light and colour between the two sources; matching the generated shots to the real footage's lighting is far easier than the reverse, so grade your decisions around whatever real material you have.