HeyGen Image to Video: the mouth pass after the picture exists (50cr)
HeyGen Image to Video wants a plate and a track. It is not how you invent the scene.
HeyGen Image to Video wants a plate and a track. It is not how you invent the scene. Catalog copy: turn your photo into a talking avatar video. Upload a selfie and choose a voice — HeyGen Avatar 4 generates realistic lip-sync video. Content type lipsync. Category audio-to-lipsync. audio is true. Requires an image. Credits 50. Durations empty. The selfie is the plate. The voice is the track. The kitchen behind them is not this row's job.
The slug is heygen-avatar4-image-to-video. The display name is HeyGen Image to Video. There is no HeyGen brand hub in the published provider list. The lipsync tool is the launcher. Best lipsync and best talking heads are the rankings.
A selfie is not a scene
requires_image is true. No photo, no job. The photo is the performer, not a mood board. Light it like a talking-head plate: face readable, shoulders in, background you can live with, because image-to-talking-avatar will not restage them in a new office.
If you needed a new office, generate or shoot the plate first on a stills row, then come here. If you needed cinematic blocking, that is text-to-video or image-to-video, not Avatar 4.
Aspects: 16:9, 5:4, 1:1, 4:5, 9:16, auto. Qualities listed: 1080p, 360p, 480p, 540p, 720p. Pick the ratio of the channel. A 1:1 selfie in a 9:16 slot is a crop fight you can avoid before you spend 50 credits.
Fifty credits starts the mouth
Price matrix is per second at 10 credits per second, with minCredits 50 and maxCredits 1200, duration-driven. The catalog display is 50. That 50 is the floor for the pass, not a 5-second text-to-video clip. Length follows the voice you chose, which is why supports_durations is empty.
Choose the voice on purpose. Features: Photo-Avatar, Custom-Image, Multi-Voice, Expressions. A silent prompt is not a path here — audio is true and the description tells you to choose a voice. The wrong voice is a 50-credit rerecord.
Do not use this row to discover whether you like the person. Discover that on a 1-credit still. Then upload the still you signed.
Voice is the second input
The pipeline is: approved photo, chosen voice (or a VO you already have, depending on how you drive the tool), then lip-sync video out. Inventing wardrobe in the prompt while the selfie shows a hoodie is how the mouth pass fights the plate.
Sync 2.0 is for footage you already shot (video-to-lipsync). LTX 2 Audio to Video is a 20-credit audio-to-lipsync cell that does not require an image. HeyGen Avatar 4 is the selfie-required, 50-credit floor, multi-voice talking-photo row.
After the file exists, captions and delivery are still downstream. This row does not publish.
FAQ
Can I run HeyGen Image to Video from text only?
No. requires_image is true. Upload a selfie or a locked portrait.
Why is duration not 5s or 10s?
The row is duration-driven off the voice, billed per second, floor 50 credits. There is no 5s/10s dropdown in supports_durations.
Is this a scene generator?
No. It turns a photo into a talking avatar. New locations belong on a stills or video row before this pass.
What does 50 credits represent?
The catalog display price and the price-matrix minimum for this lipsync job. Extra length follows per-second billing at 10 credits per second, up to the listed max of 1200.