Kling Video V3 Standard Image to Video is Kling's image-to-video model on Versely. This page is its structured prompting reference: the 5 parameters its schema actually exposes, the i2v technique that applies to it, copy-ready templates, and real prompts from production workflows that shipped on it.
Everything here is grounded in the same sources Versely's agent reads — the model's input schema, the kling family rule its prompt enhancer applies, and prompts quoted verbatim from shipped workflows. Where a line is general craft advice rather than a documented fact about Kling Video V3 Standard Image to Video, the page says so.
What Kling Video V3 Standard Image to Video wants
The exact input surface, from the same schema the Versely agent fetches with get_model_input_schema before every generation.
Variant via kie
| Parameter | What it does | Values |
|---|---|---|
prompt | Prompt — required for single-shot (multi_shots=false), optional when multi_shots=true | string |
image_urlsreq | Required for I2V — single = first frame, two = [first, last] | array (min 1, max 2) |
durationreq | Seconds — full integer range 3-15 per docs | 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15default: 5 |
aspect_ratio | Aspect ratio (auto-adapted when image_urls provided) | 16:9 · 9:16 · 1:1default: 16:9 |
multi_shotsreq | Multi-shot mode — when true, multi_prompt is required (NOT yet modeled in this schema) | booleandefault: false |
- — Both Pro and Standard share slug — differentiated by `input.mode`
- — Multi-shot composite fields (multi_prompt[], kling_elements[]) NOT yet modeled — use single-shot only for strict-mode validation
- — Verified against docs.kie.ai/market/kling/kling-3-0
Variant via fal
| Parameter | What it does | Values |
|---|---|---|
start_image_urlreq | Start frame (min 300x300, max 10MB, aspect 0.4-2.5) | string |
prompt | Required unless `multi_prompt` provided | string (max 2500) |
multi_prompt | Alternative to `prompt` | array |
duration | 3-15 seconds (string enum) | string (min 3, max 15)default: 5 |
end_image_url | Optional end frame | string |
generate_audio | Native audio | booleandefault: true |
negative_prompt | Exclusions | string (max 2500)default: blur, distort, and low quality |
The rule Versely's enhancer applies
Versely's prompt enhancer carries a per-family rule for kling models, applied automatically whenever it rewrites a prompt for Kling Video V3 Standard Image to Video. Verbatim:
“Kling video models work best with clear scene descriptions including camera movement (pan, zoom, tracking shot), subject action, and environment details. Specify motion direction and pacing.”
Technique that applies here
Image-to-video: animating a start frame; motion description relative to the input image; UGC hook pool applies
- Kling and Wan's MODEL_TIPS carry over unchanged from their text-to-video siblings, because the enhancer matches by model name, not by mode. Kling's tip ('camera movement (pan, zoom, tracking shot), subject action, and environment details... motion direction and pacing') and Wan's ('straightforward scene descriptions and style keywords') apply identically to Kling Video V3 Standard Image to Video and Wan 2.7 Image to Video as they do to each family's t2v page.
- Several i2v schemas quietly accept a second, optional 'end' image alongside the required start image — Vidu Q3's fal variant (end_image_url, described as generating a transition), Kling Video V3 Standard Image to Video's fal variant (end_image_url), LTX 2.3 Image to Video Pro (a model folded into the LTX 2 Pro page; end_image_url), and Wan 2.7 Image to Video (last_frame_url). If you supply one, write the prompt as the path between the two states, not as 'what happens next' from a single photo — you're now effectively writing a first-last-frame prompt inside an i2v call.
- General technique: keep the description to one continuous physical action sized to the fixed duration enum you picked — Hailuo only offers 6s or 10s, Kling's i2v duration runs roughly 3-15s depending on tier. Because the model has to reconcile your text against a real uploaded photo, a second competing action or a contradicting detail (different clothing, different background) fights the source image rather than animating it.
Copy-ready templates
Replace the bracketed slots; each template says when it's the right shape.
[SUBJECT] [ACTION], camera [MOVEMENT TYPE — pan / zoom / tracking shot]. Motion direction: [e.g. left-to-right / toward camera]. Pacing: [slow / steady / quick].
Use when: Kling-family i2v models — the applied family tip rewards named camera movement plus explicit motion direction and pacing.
[STATE A DESCRIPTION] transitions into [STATE B DESCRIPTION] as [WHAT CHANGES IN BETWEEN].
Use when: you've supplied both a start image and an optional end/tail image (Vidu's end_image_url, Kling Video V3's end_image_url, LTX 2.3 Image to Video Pro's end_image_url, Wan 2.7's last_frame_url) — describe the whole path, not a single continuation.
Real prompts that shipped on Kling Video V3 Standard Image to Video
Quoted verbatim from Versely's production workflow library — each one generated a scene in a shipped workflow.
Use the character reference image. Photorealistic vertical 9:16 video, authentic phone-selfie UGC style, like a real lifestyle creator casually filming herself. Soft natural morning light, slight handheld camera shake, 35mm lens look, shallow depth of field, 4K, natural film grain, realistic skin texture with visible pores and freckles, natural eye blinks, realistic lip-sync, clear natural voice, quiet room tone, no music. WOMAN — late 20s, warm sun-kissed tan skin with light freckles, hazel-green eyes, full natural brows, fresh minimal clean-girl makeup, soft nude lips. Light-brown highlighted hair in a soft loose low bun with a few face-framing strands. Wearing a cream-oatmeal ribbed knit sweater, small gold hoop earrings, thin gold necklace and ring. ENVIRONMENT — bright airy modern kitchen, warm white wall, light wood counter. Softly blurred behind her: a pale ceramic matcha bowl, a bamboo whisk on a stand, eucalyptus sprigs in a vase. Gentle morning sun from a side window. ACTION & DIALOGUE (calm, natural, NOT exaggerated): She holds her phone in selfie position, relaxed and casual, looking into the camera with a soft natural expression — calm, friendly, understated, like talking to a friend. Minimal subtle hand movement, natural small head tilts, occasional natural blink, no theatrical gestures or big smiles. REALISM: Subtle, grounded, true-to-life behavior. No over-acting, no exaggerated emotions, no fast movements. Looks like a candid real video, not AI-generated.
Shipped in Morning Matcha Routine (scene: Intro selfie).
Begin from the input frame — OTS street interview on a NYC sidewalk at golden hour. Riley stands partially visible on the left holding a black handheld mic extended toward Kai. Kai stands center-frame in burgundy hoodie and grey beanie. Camera holds steady with subtle natural handheld breath, an almost imperceptible drift — like a real iPhone or handheld DSLR shot by a content creator. 0-1.5s: Riley speaks from off-frame left in a warm friendly real American conversational voice — natural breath before the line, casual approachable energy, slight natural intonation lift on the word "wildest": "Hey — what's the wildest thing you've ever done?" Mic stays steady in frame. Kai listens, single natural blink, the corner of his mouth lifting subtly, eyes briefly flicking toward Riley. 1.5-9s: Kai answers. Slight casual shrug as he begins, eyes settle on Riley off-frame left. He delivers the line in a calm dry deadpan American conversational tone — real natural human voice with audible breath, throat catch, micro-pauses between phrases, slightly mumbled storytelling energy, real organic mouth shapes — NOT a polished voiceover read: "I broke into an abandoned subway station… lived down there for a week. Tattooed myself by flashlight." A barely-there knowing smirk lands fully on the final phrase. His left hand subtly raises the iced coffee a couple inches as he speaks, unconscious natural fidget. Background remains lively and authentic — softly blurred pedestrians continue walking past at natural unsynchronized rhythms, a yellow taxi drifts through the distant background, faint subway steam rises near the curb at right, golden-hour light flickers very subtly as people pass through it. Real urban audio bed underneath the dialogue — soft city traffic, a distant car honk, light footfall, faint conversation drift, ambient NYC murmur. CRITICAL VOICE DIRECTION — must sound 100% human, NOT AI voiceover. Both Riley and Kai must sound like real humans being filmed on the street with a hand mic. Natural breath, real throat sounds, slight verbal imperfections, real intonation variation, real human conversational cadence with micro-pauses, real emotion in the voice. Kai = slightly lower vocal register, dry deadpan Brooklyn-creative delivery with a hint of vocal fry, casual relaxed pace. Riley = warm bright friendly conversational lift, slightly higher register, no broadcast-perfect pronunciation. NO AI synthesizer cadence, NO robotic pacing, NO uniform smooth vowels, NO flat affect, NO over-enunciation, NO broadcast voiceover tone. Visible real skin texture preserved, real tattoo aging, natural micro-blinks (2-3 across the clip), natural weight shifts, slight subtle head movement during speech. Realistic handheld documentary cinematography. No robotic motion, no plastic skin, no synthetic smoothness, no over-perfect lip sync, no exaggerated theatrical gestures, no over-staged performance. Real street vlog energy — looks and sounds like a TikTok/Instagram street interview clip shot today on a real handheld setup.
Shipped in NYC Street Interview (scene: Kai — Abandoned Subway Story).
A photorealistic 25-year-old American woman sitting in a cozy podcast studio, exactly matching the reference image: brown hair in a soft low updo with loose face-framing strands, gold hoop earrings, thin gold necklace, cream ribbed tank top under an oatmeal oversized knit cardigan slipping off one shoulder, dark pants, seated on a soft light couch, a black studio microphone on a boom arm in the foreground to her right, a warm circular glowing wall light behind her, a tall green plant softly out of focus on the right. She is talking naturally and warmly into the microphone. Natural relaxed lip-sync matched to the voiceover, subtle realistic mouth and jaw movement. Expressive but calm podcast delivery: gentle confident smile, light natural head tilts and nods, soft blinking, relaxed shoulders, occasional small hand gesture near her lap. The leaves of the background plant sway very slightly. Warm ambient lighting stays steady, soft golden glow on her face. Camera is locked on a tripod with an extremely subtle slow push-in, static podcast-clip framing, vertical 9:16, shallow depth of field with creamy bokeh. Ultra photoreal skin texture, natural color grade, candid authentic expression, 4K, cinematic. Realistic human motion only, smooth and natural. Negative prompt: cartoon, anime, illustration, CGI, 3D render, plastic skin, over-smoothed, distorted face, warped mouth, bad lip-sync, extra fingers, deformed hands, asymmetric eyes, jitter, fast jerky motion, morphing, floating objects, text, watermark, logo, changing outfit, identity change.
Shipped in Podcast Clip — Versely (scene: Hook — income tease).
How the Versely agent does this automatically
You can use this page by hand, or let the agent apply the same knowledge. Four real mechanisms — no more, no less:
get_model_input_schema— before generating, the agent looks up Kling Video V3 Standard Image to Video's exact input fields, required fields, allowed values, defaults, and min/max bounds. The parameter table above is that same surface.- The prompt enhancer's family rules — 12 per-family rewrite rules, including the kling rule quoted on this page, shape how a rough prompt gets rewritten.
- The per-provider speech guide — for TTS scripts, the agent follows a provider-specific tag scheme — not relevant to this model, but it's why voiceover scripts come out marked up correctly.
expand_movie_scene— in movie flows, brief scene ideas are rewritten into detailed cinematic descriptions before generation.
Mistakes that waste generations
- Asking for a specific aspect ratio in the prompt when the schema doesn't expose an aspect_ratio param — it's silently ignored; on Vidu, Pixverse, and Wan's i2v variants the frame shape comes entirely from the source image you upload.
- Re-describing the subject or scene the source image already shows instead of focusing on the motion — wastes prompt budget and can conflict with the photo (a different outfit or background than what's actually in frame).
- Treating every 'image-to-video' model as the same shape: sending a HeyGen-style voice/talking_style prompt to Vidu, or a plain motion-description prompt to HeyGen, targets a control surface that model doesn't expose.
The long-form guide
This page is the structured reference. For the essay treatment — worked examples, failure modes, and narrative — read Kling O3 Prompting: Camera, Motion, and Reasoning.
This guide also covers
These siblings share Kling Video V3 Standard Image to Video's prompting-relevant input surface, so their prompting URLs resolve here — tier and pricing differences live on their own model pages:
Frequently asked questions
Does Kling Video V3 Standard Image to Video support negative prompts?+
Yes — the schema exposes negative_prompt (max 2500 characters). Put exclusions there instead of writing "no text, no watermark" into the main prompt.
Which aspect ratios does Kling Video V3 Standard Image to Video support?+
The aspect_ratio parameter is an enum: 16:9, 9:16, 1:1. Set the parameter — describing the frame shape in prose does nothing on its own.
How long can a Kling Video V3 Standard Image to Video generation be?+
Duration is a hard enum: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15. Write one continuous beat sized to the window you pick, not a script the model will compress.
How does the Versely agent know Kling Video V3 Standard Image to Video's parameters?+
Before generating, the agent calls its get_model_input_schema tool, which looks up the exact input fields, required fields, allowed values, defaults, and min/max bounds for the model; separately, the prompt enhancer applies the kling family rule quoted on this page to the prompt text itself. Nothing on this page is guessed — it is the same schema surface those tools read.
Does this guide also cover Kling V2.1 and Kling O1 Image to Video and others?+
Yes. Kling V2.1, Kling O1 Image to Video, Kling O3 Pro Image to Video, Kling O3 Standard Image to Video share the same prompting-relevant input surface as Kling Video V3 Standard Image to Video, so their prompting URLs redirect here instead of duplicating this page. Tier and pricing differences live on each model's own /models page.
Related prompting guides
Generate with Kling Video V3 Standard Image to Video
Kling Video V3 Standard Image to Video is live in Versely — paste a template above, or just describe what you want and let the agent map it onto the schema for you.