Guides

    HeyGen Avatar V3: the mouth pass after the picture exists (17cr)

    HeyGen Avatar V3 wants a plate and a track. It is not how you invent the scene.

    Versely Team5 min read

    HeyGen Avatar V3 wants a plate and a track. It is not how you invent the scene.

    HeyGen Avatar V3 generates digital twin videos from 700+ premade avatars, with customizable voices, expressions, and styles. Category is premade-avatars. Content type is lipsync. Audio is on. An image is not required — the plate is the avatar you pick, not a still you designed. The catalog lists 17 credits. Aspect ratios: 16:9, 5:4, 1:1, 4:5, 9:16. Qualities: 1080p, 360p, 480p, 540p, 720p. No duration ladder on the slug; the matrix is per-second with a 17-credit floor.

    The picture already exists in the roster. Your job is the line, the ratio, and the honesty that this will not design a new room.

    The twin is the plate

    700+ premade avatars is the catalog number. You are casting from a library, not generating a location. If the brief is "our founder in our actual office with our actual bottle," this is the wrong lipsync. If the brief is "a presentable talking person saying this script," V3 is the 17-credit floor for that job.

    requires_image is false because you do not bring a still. You bring a choice. Customizable voices, expressions, and styles are the knobs on that choice. They are not a set-build. Do not prompt V3 to invent a factory tour. Cast a twin. Write a line. Pick 9:16 or 16:9 from the list that is actually published.

    Audio is on because speech is the point. The track is the script or a voice you attach through the product. This is still a mouth pass: the face was already modelled. You are not "inventing a scene with dialogue." You are putting words on a premade digital twin.

    Seventeen credits is the listed floor. Billing is duration-driven in the matrix (up to 408). A long read is a long bill. Cut the script to the line you will ship before you spend the floor on throat-clearing.

    What "after the picture exists" means here

    On Wan 2.2 Speech to Video, the picture is your still and the model requires it. On HeyGen Avatar V3, the picture exists in the 700+ library. Same rule, different plate: do not use this row to design the world. Use it after you have accepted a twin as the world.

    1080p is on the quality list, down to 360p. There is no 4K token. A talking social close-up is the spec. Iterate at 720p or 540p if you are still casting. Spend 1080p on the twin and the line you would keep.

    Run it from AI lipsync, not from a text-to-video box. HeyGen has no provider hub on this site. Best lipsync and best talking heads are the job maps. Cheapest lipsync is the price map — this row is 17 credits listed, not the floor of the catalog. This page is V3's premade-twin rule.

    A twin is not a warehouse

    A custom location. A product hero that is not a talking person. A silent cinematic plate. A 20-second landscape with no one in it.

    If you need a scene, generate the scene on a video row (or shoot it), then use a video-to-lipsync or image-to-lipsync model that accepts that plate. V3 is premade-avatars. Mis-labelling it as "HeyGen, so it can do our brand film" is how 17 credits buys a twin in a generic void saying a launch line that needed a warehouse.

    Captions still after. A premade mouth is not a transcript. Burn the script once the take is the take.

    Is a library twin allowed?

    If no — likeness, founder, specific wardrobe — stop. You needed a still of the real person and a different lipsync row.

    If yes, pick the twin, pick 9:16 or 16:9, write a short line, spend the 17-credit floor. Do not first ask V3 to also be the set designer.

    If you have no script and no voice, you do not have a track. This model is lipsync-shaped. Cast, then speak.

    FAQ

    Do I upload my own photo to HeyGen Avatar V3?

    Not as a requirement. The catalog marks requires_image false and files the row under premade-avatars, 700+ twins. If you have a specific face that must be yours, that is a different lipsync (image-to-lipsync) with a plate you supply. V3 is the library.

    Why 17 credits with no duration list?

    17 is the listed figure and the matrix floor for a per-second lipsync job. There is no 5s/10s menu on the slug. Length follows the take. Cut the script first. A long twin read is how the floor stops being the bill.

    Can this replace a text-to-video scene generate?

    No. It will not invent the warehouse, the pack orbit, or the crowd. It will put customizable voice and expression on a premade avatar at up to 1080p. Scene first (or twin-as-scene), mouth second.

    Is native audio a soundtrack?

    Audio is the talking. Treat it as speech, not as a music bed. If you need a score, add it after. If you need a silent plate, you are on the wrong row.