Two files go in. The base is the plate the audience came for; the overlay is a second take, scaled and parked in a corner. Neither file is invented by the composite. If you do not have both, this is the wrong job.
A generate that tries to put a presenter and a product in one prompt is a different model call with a different identity risk. Captions and text overlay sit on top of one file; they do not introduce a second video. A recurring series that always uses the same inset is a format decision, not a new generate.
The editing task is add picture-in-picture to video. The agent can overlay a talking head on a video you already have. Premade avatars and lipsync create the inset; they do not place it.
In practice
- Finish the base plate first; the inset is a later composite, not a second prompt on the same generate.
- Keep the inset off the burned-in caption row and off platform UI chrome.
- If the brief is one talking presenter and no B-roll, skip PiP and shoot or generate the presenter as the whole frame.
The mistake to avoid
Asking one generate to invent both the product shot and the host, then calling the muddy result picture-in-picture. PiP is a composite of two files you already have.
Where you will run into it
- Add Picture-in-Picture to a Video — Reaction footage, corner-mounted on your main clip.
- AI Video Editor — Upload the clip. Say what's wrong with it.
Related terms
B-roll
B-roll is the supporting footage cut over narration or an interview — everything on screen that is not the person doing the talking.
Burned-in captions
Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.
Premade avatar
A premade avatar is a stock presenter from a ready-made roster you pick, rather than a face you supply or a mouth you lipsync onto footage you already shot.
Lipsync
Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.