Guides

    Kling 3.0 Prompting Guide

    Prompt Kling 3.0 and Kling 3 Turbo as a shot brief: subject, motion, camera, multi-shot, motion control. Run the 22-credit Turbo rows on Versely.

    Versely Team11 min read

    Kling 3.0 prompting is a shot brief, not a mood board caption. Name who is in frame, what moves, which shot type and camera move own the clip, and (when you need cuts) which beats land in which numbered shot. On Versely the speed path is Kling 3 Turbo Text to Video and its image-to-video twin at 22 credits in the catalog snapshot (3-15s, up to 1080p). Iterate the brief instead of restacking "cinematic 8K" adjectives.

    This guide is the technique layer for the Kling 3.0 / Kling 3 Turbo family. It is deliberately distinct from the Kling O3 prompting guide, which owns reference locks and Omni-style identity stacks. For Turbo's silent contract see how to prompt when Kling 3 Turbo is silent. Sibling technique hubs: Seedance 2.5 prompting and Wan 3.0 prompting.

    The shot brief stack

    Write motion over time. A still description gives the model nothing to animate.

    1. Subject - who or what must stay readable
    2. Action - one verb chain the camera can hold
    3. Setting + light - place, time of day, key direction
    4. Camera - shot size plus one primary move
    5. Format - duration and aspect (3-15s; 16:9, 9:16, or 1:1)
    6. Continuity - face, wardrobe, product traits that must not reset
    7. Audio - only when the row supports native sound (see below)
    8. Constraints - exclusions in the prompt text

    Weak: Cinematic shot of a woman in a city at night.

    Stronger: Handheld medium tracking shot following a woman in a dark coat through rain-slicked streets. She turns a corner, slows, and checks her phone. Cool neon side light, shallow depth of field. 6 seconds, 9:16.

    Worked prompt 1 (text-to-video, single shot)

    FORMAT: 9:16, 6 seconds, photoreal commercial.
    SUBJECT + ACTION: A frosted citrus soda bottle on wet slate. A hand lifts it and pours a steady stream into a tall glass of ice until foam rises and settles.
    CAMERA: Close-up, slow push-in locked on bottle and glass. No orbit.
    LIGHT: Bright daylight from upper left, crisp highlights on the wet slate.
    CONTINUITY: Same bottle label, same glass, same ice level progression.
    CONSTRAINTS: No text, no subtitles, no warped hands, no extra bottles.
    

    Camera language that Kling actually reads

    "Cinematic" is not a move. Name the shot and one move, then stop.

    Useful vocabulary: locked-off, slow push-in / dolly-in, pull-back, tracking / follow, orbit / arc, pan, tilt, rack focus, handheld, low angle, eye level, medium close-up, wide.

    Tie the move to an on-screen event when timing matters: As the food leaves the pan, the camera orbits the chef's face, then settles. Stacking "dolly, orbit, whip, crane" in one sentence without triggers makes the model average everything into mush.

    Separate subject motion from camera motion in different clauses so neither flattens the other.

    Multi-shot: one action per numbered beat

    On Kling VIDEO 3.0 / Turbo storyboard modes you can describe up to about six shots in one generate. Each shot gets a number, a duration in seconds, and its own beat. Shot durations should add up to the clip length. Keep one primary action per shot; chain shots instead of cramming three ads into one second.

    shot 1, 3s, wide low-angle of a cyclist cresting a mountain switchback at dawn, cool side light;
    shot 2, 2s, medium tracking shot on her face, breath visible in the cold;
    shot 3, 2s, close-up of pedals turning, gravel spray under the tires
    

    If Multi-Shot / Custom Multi-Shot is available on the provider UI, turn it on when you want cuts. If it is off, write a single continuous take and do not fake edits with semicolon spam the model will ignore.

    Worked prompt 2 (multi-shot brief)

    FORMAT: 16:9, 8 seconds total.
    shot 1, 3s, locked-off wide of a ceramic pour-over dripper on a wood counter at sunrise; steam rising;
    shot 2, 3s, medium close-up of a barista in a navy apron pouring slow spirals of hot water;
    shot 3, 2s, slow push to the coffee bed blooming in the filter
    CONTINUITY: Same navy apron with cream patch, same counter wood grain, same morning window light.
    CONSTRAINTS: No subtitles, no logo invent, no extra people.
    

    Motion control vs plain T2V / I2V

    These are different contracts. Do not mix them in one mental model.

    Mode What you give What the prompt does
    Text-to-video Text only Builds subject, scene, motion, and camera from language
    Image-to-video Still + text Still owns look; prompt owns what happens next and how the camera moves
    Motion Control Character image + motion reference video (+ optional prompt) Transfers performance from the video onto the character

    Matches Video (character orientation follows the reference video): best for complex full-body choreography. Pose, facing, and often camera follow the motion clip. Use a clear full-body or upper-body driver and a matching-proportion still.

    Matches Image (orientation follows the still): keeps the character facing the way the photo faces. Motion and expression still come from the video, but camera is largely prompted (Zoom In, Zoom Out, Camera Up, Camera Down, Fixed Position on Kling's own guide). Cap tends to be shorter than Matches Video.

    Prompt motion control for background, wardrobe notes you must keep, and (in Matches Image) the camera treatment. Do not rewrite the entire dance in prose when the video already owns the timing.

    Worked prompt 3 (image-to-video: direct the still)

    From the image: she lowers the coffee cup to the saucer, then turns toward camera and smiles. Soft curtain stirs behind her.
    CAMERA: Nearly locked medium close-up with a slight slow push after the smile.
    CONSTRAINTS: Keep her face, hair, and sweater from the first frame. No new room. No subtitles. No warped hands.
    

    Do not re-describe the photo ("a woman in a red sweater sits by a window…"). The pixels already answered that. Animate what is physically reachable from the pose; add one environmental motion (curtain, steam, passing traffic) so the take reads as video rather than a warped still.

    For heavy multi-reference identity stacks, switch to O3 / Omni workflows and use the O3 prompting guide. On plain 3.0 I2V, a clean first frame plus a continuity clause usually beats five conflicting refs.

    Identity hold without fighting the row

    • Prefer a clean, well-lit subject still for I2V: face readable, minimal clutter, no half-occluded limbs.
    • Restate two visible traits at risk points (after a turn, after a cut): Same short black bob, same green wool coat throughout.
    • On VIDEO 3.0 element binding (when the provider exposes it), bind the subject so zoom/pan/tilt do not erase the face. Do not also ask the prompt to invent a second wardrobe.
    • If the face drifts, fix the reference still before rewriting the whole brief.

    Audio: native on 3.0, silent on Versely Turbo

    Kling VIDEO 3.0 can generate native audiovisual output when the row supports it: named speakers, quoted lines, accents, and layered ambience. Write dialogue as Name (tone): "exact line" and keep one speaker clear per beat.

    On Versely, Kling 3 Turbo text-to-video and image-to-video are silent plates. Quotes will not become a stem. You may get a mouth that mimes. You will not get a read. Write motion; score later with music/TTS, burn captions after, or pick a native-audio row (Seedance, Veo, Grok Imagine Video, Kling O3 Pro, and similar). Details: how to prompt when Kling 3 Turbo is silent.

    Failure modes

    Failure Likely cause Fix in the brief
    Pretty still, no motion Caption without verbs or timing Write what moves over the seconds
    Camera feels random "Cinematic" with no shot or move Name shot size + one move
    Mushy middle Three actions in one shot One action per shot; use multi-shot
    Chewing face, no voice Dialogue on silent Turbo Motion-only brief; TTS/native-audio row
    Warped subject on I2V Action unreachable from the pose Bridging motion; one env. motion; cleaner still
    Face reset after a turn Continuity not restated Re-anchor two traits at the risk beat
    Motion control jitter Still vs driver mismatch Match proportions/angles; try Matches Video vs Matches Image deliberately
    On-screen text wrong Asking the model to typeset Burn captions after the take

    Skip the dialect: brief the Versely agent

    You can learn this stack. On Versely you can also skip memorizing every Kling dialect and brief the agent instead of writing prompts: content type, photos or references, the edits you will accept, target platforms, and a budget ceiling. Name the row when it matters ("use Kling 3 Turbo Image to Video" or "silent plate, score later"). The agent plans the job; you approve the plan instead of babysitting clause order.

    A hand-tuned shot brief still wins for one hero take. A job brief wins for variants, captions, and a post path with a spend cap.

    After the take: edit, post, collections

    When the Kling pass is close enough:

    1. Light trim if a beat overruns. Do not ask the model to be CapCut.
    2. Burn mute-proof captions with /tools/ai-caption-generator or the on-device /free-tools/burn-captions path. Especially useful on silent Turbo plates.
    3. Upload or schedule to the social accounts Versely already connects for your workspace. If a network is not connected, export the mp4 and post from the native app.
    4. Save keepers into a Versely collection so the next brief reuses the same stills, winning camera lines, and multi-shot templates.

    Honest limit: board separate generates when you need separate setups, then stitch. Multi-shot is for related coverage in one clip, not a feature film.

    FAQ

    What is the best Kling 3.0 prompt formula?

    Subject and one action, setting and light, shot size plus one camera move, format, continuity, then constraints. Add numbered multi-shot beats when you need cuts. Add Audio only on rows that actually generate sound.

    How is Kling 3.0 / Turbo different from Kling O3?

    3.0 / Turbo are prompt-led motion and storyboard rows (plus motion control when available). O3 / Omni is the reference-driven path for locked identity across refs and edits. Use this guide for 3.0/Turbo; use the O3 prompting guide for Omni stacks.

    Does Kling 3 Turbo on Versely generate audio?

    No. The Turbo catalog rows are silent. Write motion, then score or caption, or switch to a native-audio model for talking plates.

    How do I write multi-shot prompts?

    shot 1, 3s, …; shot 2, 2s, … with durations that sum to the clip. One primary action per shot. Up to about six shots when the mode is on.

    When should I use Motion Control instead of I2V?

    When you need a specific performance or trajectory from a reference video transferred onto a character still. Use Matches Video for full-body follow; Matches Image when facing must stay with the photo and camera is prompted.

    How much do the Kling 3 Turbo rows cost on Versely?

    Catalog snapshot: 22 credits for text-to-video and image-to-video (billed with length and quality in the composer). See /models/kling-3-turbo-text-to-video. Do not assume a discount from this post; trust the live composer total before you send.

    Where do I run this on Versely?

    Open the AI video generator with Kling 3 Turbo selected. Use this page for the brief; use the silent-Turbo how-to when you need the audio failover tree.

    Takeaway

    Prompt Kling 3.0 and Kling 3 Turbo like a short shoot: one subject, one action, one camera move, and numbered shots when you need cuts. Treat Motion Control as a transfer job, not a prettier T2V. Keep O3 for reference locks. On Versely Turbo, plan a silent plate and score or caption afterward. Run the 22-credit rows, or hand the agent a job brief when you would rather not memorize the dialect. Then caption, post where the product supports it, and file keepers in a collection so the next take starts warmer.