How to prompt when Kling 3 Turbo is silent
On Versely, Kling 3 Turbo T2V is audio false, 3–15s, 1080p, 12 credits discounted. Write motion, not dialogue. If you need a read, generate the picture here and attach TTS or pick a native-audio row.
On Versely, Kling 3 Turbo T2V is audio false, 3–15s, 1080p, 12 credits discounted. Write motion, not dialogue. If you need a read, generate the picture here and attach TTS or pick a native-audio row.
That is the whole prompt rule. Kling 3 Turbo Text to Video is a fast silent plate. Quoted lines do not become a stem. “Epic soundtrack” does not become a bed. A mouth in frame will not deliver the joke you typed in quotes. Prompt the camera and the body. Score somewhere else.
Kuaishou's own Turbo write-ups talk about bundled audio at the source. Versely's published row does not. Kling 3 Turbo: Fast Native-Audio Clips for Daily Posting is the asterisk on that split. This page is the prompt you write given audio: false. Do not spend a generation discovering the file is mute.
Silent is the catalog, not a failed take
audio is false on this slug. Durations are 3s through 15s, every second. Ceiling is 1080p. The catalog figure is 12 credits, discounted. No still is required. Aspects in production use are 16:9, 9:16, and 1:1. That is the object.
A silent flag is a contract. Native-audio rows — Seedance 2.5 (4–30s, 720p, 47 credits, audio true), Veo 3.1 (4 / 6 / 8s, 4K, 40 credits, audio true), Grok Imagine Video (6 / 10 / 15 / 20 / 30s, 720p, 7 credits, audio true), Kling O3 Pro Text to Video (3–15s, 70 credits, audio true) — speak in the same pass as the picture. Turbo will not. A Veo-shaped dialogue brief here is a chewing face and a silent export.
Open the row from the AI video generator. Do not open it from a habit that says “Kling always talks now.” Other Kling slugs have other flags. Read Turbo.
Write motion, not dialogue
Strip every token that assumes a soundtrack.
Cut these: quoted speech, “she says,” “voiceover,” “narrator,” “lip-sync to,” “whoosh,” “boom,” “epic score,” “diegetic radio,” “the line is…”. They waste prompt budget and they cue a mouth the model cannot voice.
Keep these: subject, verb, camera, light, lens, duration you will actually buy.
A usable shape:
[who/what] + [what moves] + [how the camera moves] + [light] + [lens] + [how long]
Weak: A barista says "try the new cold foam" and the espresso machine whooshes.
Strong: A barista in a linen apron pulls a double shot on a chrome group head, crema rising in the cup, slow push-in from chest height over 5 seconds, window light camera-left, 35mm, shallow depth of field, no text, no logos invented.
The second prompt is a shot. The first is a radio play. Turbo can attempt the shot. It cannot attempt the radio play.
Motion language that actually changes the file: slow dolly in, locked-off wide, handheld follow at hip height, quarter-turn, fabric settling, steam drifting camera-right. If you bought 5s, write a 5-second action. A 15-second story inside a 5-second clip smears.
Do not ask Turbo to typeset a caption. Burn type after. This page is text-to-video. A pack shot that must match an existing still is a different job.
Start at 3s, 4s, or 5s unless the beat is longer. A 15-second silent miss is a long miss. Turbo exists to iterate. Use the cheap end of the slider to learn whether the motion is even Kling-shaped.
Attach a read, or change the row
After picture lock you have three honest sound paths. Pick one before you generate, not at export.
1. The mouth is not the shot. Product turn, hands, landscape, machine, pour, walk-away. Generate silent on Turbo. Lay text to speech or a music bed on a plate where we do not see lips. Captions carry the mute-autoplay version. Off-screen VO on a silent plate is a valid pipeline.
2. The mouth is the shot, and any plausible voice will do. Do not use Turbo. Write the line on a native-audio row: Seedance 2.5 for a long single pass, Veo 3.1 for a short 4K talking plate, Grok Imagine Video for an Imagine-looking 10s or 30s, Kling O3 Pro to stay in-family with audio on. Native audio versus TTS is the tree. Lips in frame want native audio first.
3. The mouth is the shot, and the voice is a locked brand read. Turbo still cannot speak. Use a talking native-audio take, or a silent plate without a chewing mouth plus lipsync. Do not lay a cloned founder voice on a Turbo face that was never directed to speak.
The failure mode is mixing 1 and 2 on one mouth: silent Turbo plus TTS on visible lips, no lipsync pass. That is a talking-head that is not talking.
A silent brief you can rerun
Write the prompt as a shot list even when you only buy one clip. One verb per beat. No soundtrack nouns.
5s, 9:16, 1080p. A stainless pour-over kettle tilts, water spirals onto a coffee bed in a glass dripper, steam in sidelight, slow locked-off close-up, 50mm, no people, no text, no logos. Continuous pour for the full five seconds.
That prompt has nothing for a silent model to ignore. Rerun it at 3s if the pour smears; rerun at 8s if you need more of the spiral. Change the verb, not the speech.
When you need a hook line, write it for captions or for TTS, off-screen:
Caption / VO (not in the Kling prompt): “This grind changed the cup.”
The line never enters the Turbo box. The box only gets the kettle.
If five silent variants later you still wish it talked, you did not have a Turbo job. You had a Seedance or Veo job and you spent 12-credit iterations proving it. That is cheaper than a talking miss on the wrong row — as long as you stop.
Prompt motion. Score later, or leave. The row is silent on purpose.
FAQ
Does quoting a line in the prompt make Kling 3 Turbo speak?
No. On Versely the T2V row is audio false. Quotes do not become a stem. You may get a mouth that mimes. You will not get a read. Write the motion; put the line in TTS, captions, or a native-audio row.
Should I use Turbo for a talking product demo?
Not if we see the mouth. Turbo is the silent motion pass — pours, turns, hands, walk-ins. For in-shot speech pick Seedance 2.5, Veo 3.1, Grok Imagine Video, or Kling O3 Pro, where audio is actually on. For a locked brand voice, plan TTS plus a plate without chewing lips, or a lipsync pass.
What duration and resolution should I put in the prompt?
Durations the catalog sells: 3s–15s. Resolution ceiling: 1080p. Name the duration you will buy (“over 5 seconds”) so the action fits. Do not write 4K or 21:9 into a Turbo prompt and expect the row to grow a tier.
Can I generate silent here and slap TTS on a face?
Only if we do not see the mouth, or if you budget a lipsync repair. TTS on a silent chewing face is two products fighting. Off-screen VO on a product plate is the clean Turbo-plus-speech path. In-shot speech is a different row.