Kling Lipsync: the mouth pass after the picture exists (4K, 7cr)
Kling Lipsync wants a plate and a track. It is not how you invent the scene.
Kling Lipsync wants a plate and a track. It is not how you invent the scene.
Kling Lipsync is Kling's dedicated lipsync row: video-to-lipsync, audio on, 4K ceiling, 7 credits. The catalog description is the boundary: it syncs a speaker's mouth in an existing video to new audio, built for avatar animation and translated or re-voiced talking-head clips. Features listed: Accurate, Expressive, Versatile. There is no duration menu because length follows the plate and the track, not a 5s/10s picker. This is not text-to-video.
The picture has to exist first
You do not prompt a set, a wardrobe, or a camera move on this row. You bring footage whose mouth is visible, then you bring the audio you want that mouth to match. The AI lipsync tool is the door. Inventing the scene is text-to-video or a talking generate. Re-voicing a take you already have is this.
A still photo is the wrong mental model. Category is video-to-lipsync, not image-to-lipsync. If all you have is a headshot, pick a row that accepts a still. If you have a talking-head plate and a new VO, stay here.
4K, 7 credits, one job
Seven credits is the catalog price. Billing on the record is duration-driven; the display price is 7. Max output resolution is 4K. Use that when the delivery is a re-voiced hero take, not when you still owe the hero take.
The failure mode is generating a silent scene "so we have something to lipsync." That is two models, two briefs, and a plate that was never directed for speech: profile angles, hands on the jaw, mouth in shadow. Shoot or generate the plate as if someone is speaking. Then run Kling Lipsync. Lipsync a video is the editing-task map.
Plate plus track, then stop
Audio is on because the job is audio. The new track is the performance. The existing video is the body. Do not ask this row to change the room, the cut, or the product. Translation and re-voice are the listed use. A new scene is not.
Kling's other catalog rows will invent picture. This one will not. The Kling hub is the rest of the family. This page is only the mouth pass: existing video, new audio, 4K, 7 credits.
If the cut is off-screen VO, skip lipsync. Lay the track. If the cut is a premade avatar rather than your footage, that is a premade-avatar row, not this one.
Do not invent the scene here
The test is whether you could ship the picture muted. If the muted picture is not the scene, you are on the wrong row. Kling Lipsync will not save a plate that is the wrong shot. It will sync a mouth on the shot you already made.
FAQ
Can Kling Lipsync generate the talking-head scene from a prompt?
No. It syncs a speaker's mouth in an existing video to new audio. Invent the scene on a video row, then come here.
Do I start from a still?
This category is video-to-lipsync. Bring footage. A still-to-talking job is a different family.
What does 4K refer to?
Max output resolution is 4K. The 7 credits are the catalog price for this lipsync generate, not a new 4K scene.
Is this the same as cloning a voice?
No. Cloning writes a voice file. This row takes a track you already have (cloned or recorded) and drives the mouth on a plate you already have.