Wan 2.2 Speech to Video: the mouth pass after the picture exists (720p, 50cr)
Wan 2.2 Speech to Video wants a plate and a track. It is not how you invent the scene.
Wan 2.2 Speech to Video wants a plate and a track. It is not how you invent the scene.
Wan 2.2 Speech to Video is a lipsync row. Categories are audio-to-lipsync and image-to-lipsync. An image is required. Audio is on. Max output resolution is 720p, with 480p, 580p, and 720p listed. The catalog lists 50 credits. The description: high-quality video from a static image plus audio, with facial expressions, body movement, and camera work. There is no duration ladder on the row — length follows the track you bring.
That is a mouth pass. The room, the wardrobe, the product in hand: those are already in the still. If they are not, you are in the wrong office.
Plate first
requires_image is true. No still, no generate. The still is the set. Wan 2.2 will perform on it — face, body, a camera move — but it will not design the set from a paragraph. If the background is a messy bedroom and you needed a studio, recrop or replace the still. Do not prompt the lipsync model to "make it premium."
Lock the still on an image row. Approve the face at the crop you will actually ship (head-and-shoulders versus waist-up: this model lists body movement, so the crop is a performance choice). Then bring the track.
The track is the other half. Audio-to-lipsync means you already have speech. Write it with a TTS row or record it. Do not type the line into a video prompt and hope Wan 2.2 hears it. It hears a file.
Fifty credits is the listed floor for this job. It is not a 4-credit voiceover. You are paying for a talking picture at up to 720p. Bring a plate worth that.
720p talking is the spec
The ceiling is 720p. There is no 1080p or 4K on this output list. If the master needs 4K talking, this is the wrong Wan, or the wrong family. If the job is a social talking close-up, 720p is the listed spec — stop apologising for it and crop for the face.
No duration steps in the catalog. A 6-second line makes a 6-second clip; a two-minute monologue makes a two-minute bill. Cut the audio to the line you mean before you spend 50 credits on throat-clearing. The matrix on this row is duration-driven. Long takes are how a "cheap talking test" becomes the most expensive still you ever uploaded.
Camera work and body movement are in the description. That is more than a mouth mask. Give the still space to move — shoulders in frame if you want shoulders — and do not expect a locked-off newsreader if you fed it a full-body fashion plate. Conversely, a tight crop will not grow legs.
Run it from AI lipsync, not from a text-to-video box. The Wan roster has generators that do invent scenes. This slug is the speech-to-video still. Best Wan model is the family map. Best lipsync is the job map.
720p will not grow a set
A location. A second character. A product that is not in the still. A 4K plate. A soundtrack that is not the speech file — audio is the line, not a score.
If you need the scene to exist first, generate or shoot the still (or a silent clip on a different row), then come back. Wan 2.2 Speech to Video is the pass after the picture exists. Skipping that order is how 50 credits buys a talking mess with a great read.
Captions still go on afterwards. A moving mouth is not a transcript. Burn the line in the caption tool once the take is the take.
Two files, then 50 credits
Hold the still. Play the audio with your eyes closed. If either file would embarrass you, do not combine them.
If both files are good, this is the 50-credit, 720p row. Do not first "see if the model can also change the jacket." That is an image edit, and it belongs before this pass.
If you have no still, you wanted a premade avatar or a talking generate, not Wan 2.2 Speech to Video.
FAQ
Can Wan 2.2 Speech to Video start from text only?
No. The catalog requires an image, and the categories are audio-to-lipsync and image-to-lipsync. A prompt without a plate is a different Wan. Bring a still and a track.
Why 720p and not 1080p?
That is the listed max on this row: 480p, 580p, 720p. It is a talking close-up spec, not a cinematic master spec. If you need more pixels, change rows after you admit this one is a mouth pass.
Does the 50-credit figure cover any length?
50 credits is the listed figure and the matrix floor. Billing is duration-driven. A longer track costs more. Cut the audio to the line you will ship before you generate.
Should I generate the scene and the speech in one click?
Not on this model. Scene is the still. Speech is the file. One click that tries to do both is how you get a new face saying a line you already recorded for someone else. Picture first, mouth second.