Generation modes

    Picture-in-picture

    Also called PiP, Talking-head overlay.

    Picture-in-picture is compositing a smaller video — usually a talking head or reaction — over a base clip, rather than generating a new clip that already contains both.

    Two files go in. The base is the plate the audience came for; the overlay is a second take, scaled and parked in a corner. Neither file is invented by the composite. If you do not have both, this is the wrong job.

    A generate that tries to put a presenter and a product in one prompt is a different model call with a different identity risk. Captions and text overlay sit on top of one file; they do not introduce a second video. A recurring series that always uses the same inset is a format decision, not a new generate.

    The editing task is add picture-in-picture to video. The agent can overlay a talking head on a video you already have. Premade avatars and lipsync create the inset; they do not place it.

    In practice

    • Finish the base plate first; the inset is a later composite, not a second prompt on the same generate.
    • Keep the inset off the burned-in caption row and off platform UI chrome.
    • If the brief is one talking presenter and no B-roll, skip PiP and shoot or generate the presenter as the whole frame.

    The mistake to avoid

    Asking one generate to invent both the product shot and the host, then calling the muddy result picture-in-picture. PiP is a composite of two files you already have.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.