Kling Avatar Pro: the mouth pass after the picture exists (4K, 58cr)
Kling Avatar Pro wants a plate and a track. It is not how you invent the scene.
Kling Avatar Pro wants a plate and a track. It is not how you invent the scene.
Kling Avatar Pro is a professional-quality talking avatar with advanced lip synchronization. Content type is lipsync. Categories: audio-to-lipsync and image-to-lipsync. Requires an image. audio is true — the output is a talking picture, not a silent still. Max output: 4K. Display price: 58 credits. Billing is per second (min 58, max 1380). There are no duration chips on the row because length follows the track you bring, not a 5s/10s video slider.
Plate plus track, then 4K mouth
Two inputs. The still is who we are looking at. The audio is what they say. Kling Avatar Pro's job is to put those together at 4K with advanced lip sync. If you do not have the still, you are not in this row yet. If you do not have the track, you are not in this row yet.
Inventing the scene — the room, the wardrobe, the product, the first likeness — is an image generate. Inventing the line is text-to-speech. This model is the last pass. The AI lipsync tool is the door. Lipsync is the term. Ranked lists: best lipsync model and best model for talking heads.
The Kling roster is large. Avatar Pro is the professional talking-avatar line, not the text-to-video line. Do not send a scene prompt here and expect a wide shot. There is no "the camera dollies" on a lipsync row. There is a face and a line.
Fifty-eight credits after the picture exists
58 credits is the display price, the floor of a per-second bill that can run to 1380. That number is why you do not use Avatar Pro as a sketch pad. Sketch the still on a text-to-image row. Sketch the line on AI text-to-speech. Approve both. Then spend 58 credits (or more, if the track is long) on the mouth.
4K is the output ceiling. Bring a plate that can hold 4K. A small, soft still will be a 4K view of softness with a perfect jaw. The mouth pass does not invent resolution the picture never had.
requires_image is true. A text-only "talking host in a studio" prompt is the wrong label. That is how you burn the 58-credit floor on a generate that could not start.
It will not invent the scene
Background, wardrobe, the product in hand — if they are not in the plate, they are not in the take. Audio-to-lipsync and image-to-lipsync mean: drive the face from the track, using the picture you supplied. They do not mean "build me a commercial."
If the scene is wrong, edit the still. If the line is wrong, recut the wav. Then run Avatar Pro again. Do not prompt the lipsync model to "make the room warmer." That is not a lever it has.
The longer stack — clone, TTS, lipsync — is dubbing, lipsync, and voice cloning. This page is the routing rule for the Pro avatar row: picture exists, track exists, then 4K mouth at 58 credits.
The mouth pass is last on purpose
Order is the product. Still → speech → Avatar Pro. Reverse that and you either have a talking head with the wrong face or a 4K lipsync of a scratch line you will throw away. 58 credits is not a scratch price.
If you only needed VO under b-roll, you never needed this row. A TTS file is enough. Avatar Pro is for when we see the mouth.
FAQ
Can Kling Avatar Pro generate the host from a text prompt?
No. It requires an image. Categories are audio-to-lipsync and image-to-lipsync. Invent the still on an image row, then come here.
Does the 58-credit price include the voice?
No. 58 credits is this lipsync pass (per-second billing, floor 58). The track is a separate speech or clone generate. The plate is a separate still.
What resolution does the talking clip come back at?
Max output is 4K. Bring a plate that can hold it. Avatar Pro is the mouth, not an upscaler for a tiny still.
When should I skip Avatar Pro?
When the cut does not show a mouth. Lay a TTS wav under the picture instead. This row is the talking pass after the picture and the track already exist.