Comparisons

    Wan 2.7 vs Kling O3: Reference-to-Video Compared

    Wan 2.7 vs Kling O3 reference-to-video compared: character and product consistency, voice cloning, camera control, and which model holds identity.

    Versely Team6 min read

    Reference-to-video is the capability that turned AI video from a toy into a production tool. Text-to-video gives you a character; reference-to-video gives you your character — the same face, the same product, the same mascot, held across every scene of a campaign. It's the difference between generating clips and building a content franchise.

    Two models currently anchor the serious end of this category: Wan 2.7, the open-weight heavyweight whose reference mode ships with a voice-clone party trick, and Kling O3 Standard, the reasoning-enhanced Kling generation applied to reference work. I tested both on the two jobs reference-to-video actually gets hired for — a recurring human character across five scenes, and a physical product held accurately across an ad's worth of shots — plus the edge cases that expose the seams. Same reference images to both, multiple takes per scene.

    Designer's desk with reference images and sketches spread out

    What "reference" means to each model

    Both models accept reference images and generate scenes containing the referenced subject, but they weight the reference differently, and the difference drives everything downstream.

    Wan 2.7 treats the reference as near-canonical. Faces come back with high fidelity — the character in scene five is convincingly the person in your reference photo, down to secondary features like ear shape and hairline. The cost of that fidelity: Wan is more conservative about transforming the subject. Ask for your character in dramatically different lighting or stylization and Wan sometimes drags reference-image lighting along with the face.

    Kling O3 treats the reference as a strong prior it reasons over. The subject integrates more naturally into whatever scene you describe — lighting, angle, and mood adapt fluidly — but identity precision sits a notch below Wan's. The character is clearly the same person; a portrait-mode close-up comparison against the reference reveals mild feature drift, especially after unusual angles.

    The one-line version: Wan holds the face; Kling holds the scene.

    Test one: recurring character across five scenes

    Brief: same woman across a morning-routine sequence — kitchen, street, office, gym, evening couch. Different lighting and wardrobe per scene.

    Wan 2.7 delivered the more recognizable character. Lining up all five scenes, Wan's set looks like five photos of one actress. Kling's set looks like one actress in scenes one through three and possibly her sister in the gym scene, where the low-angle shot pushed features furthest from the reference pose.

    But Kling's scenes were individually better directed — its camera control produced a genuinely nice push-in on the office scene, and its prompt reasoning nailed "she checks her watch and visibly decides to skip the queue" as a readable acted beat, something Wan rendered more literally and less legibly. For a fuller treatment of keeping characters stable across whole projects, the character consistency fallback-chain guide covers the multi-model strategies that build on either.

    Test two: product accuracy

    Brief: a distinctive supplement bottle (custom label, unusual cap) across four ad shots.

    Product work punishes drift harder than face work — a slightly-off human face still reads as human, but a slightly-off label reads as a counterfeit. Wan 2.7 won this test outright: label text and geometry survived at high accuracy in most takes. Kling O3 kept form and color reliably but re-typeset label fine print in about half the takes, which for a real brand is disqualifying without a comp-in-the-real-label pass in post.

    If products are your main reference workload, Wan is the default. Kling's counterplay is shots where the product is secondary — held, in-motion, mid-ground — where its scene integration looks more natural and label precision matters less.

    The voice-clone wrinkle

    Wan 2.7's reference mode carries a capability Kling doesn't touch: voice cloning in the generation pipeline. Pair the visual reference with a voice sample and your recurring character speaks in a consistent, chosen voice across scenes. For character-driven serial content — a mascot series, a recurring UGC presenter — this collapses what would otherwise be a generate-then-TTS-then-lipsync stack into one pass. It's the single biggest workflow advantage either model owns, and how Wan's stack compares against closed rivals more broadly is covered in the Wan 2.7 vs closed-source comparison.

    The scorecard

    Dimension Wan 2.7 Kling O3 Standard
    Facial identity fidelity Winner Mild drift at extreme angles
    Product/label accuracy Winner Re-typesets fine print ~half the time
    Scene integration (lighting/mood adaptation) Can drag reference lighting Winner
    Camera control Basic Winner
    Acted beats / prompt reasoning Literal Winner
    Voice consistency for characters Winner (voice clone in pipeline) Not offered
    Stylized/transformed renditions of subject Conservative Winner

    Routing recommendation

    • Product ads, brand mascots, anything where the subject must be exactly right: Wan 2.7.
    • Character-led serial content with dialogue: Wan 2.7, on the strength of the voice-clone pipeline.
    • Narrative scenes where direction quality leads and identity tolerance is looser: Kling O3.
    • The hybrid worth knowing: establish identity-critical close-ups with Wan, shoot the wider directed coverage with Kling — at feed distance, mid-ground drift is invisible, and you get Kling's camera work where it counts. Both models sit in the same picker, so the hybrid costs nothing but attention. One practical tip either way: shoot your reference images clean — neutral lighting, plain background, subject filling the frame — because every flaw in the reference gets faithfully reproduced by Wan and creatively reinterpreted by Kling, and neither outcome is what you wanted.

    FAQ

    Which is better for character consistency, Wan 2.7 or Kling O3?

    Wan 2.7 holds facial identity more faithfully across scenes — five scenes look like five photos of one person. Kling O3 integrates the character into scenes more naturally but shows mild feature drift at unusual angles. For strict identity, pick Wan.

    Which model should I use for product ads?

    Wan 2.7. Label text and product geometry survive generation at meaningfully higher accuracy; Kling O3 re-typesets fine label print often enough that product-hero work would need post cleanup.

    What is Wan 2.7's voice cloning in reference-to-video?

    Alongside visual references, Wan 2.7 can take a voice sample so your recurring character speaks in a consistent cloned voice across generations — replacing a separate TTS-plus-lipsync pipeline for serial character content. Kling O3 has no equivalent.

    Where does Kling O3 beat Wan 2.7?

    Direction: camera control, acted story beats, and adapting the referenced subject to each scene's lighting and mood. When the shot's craft matters more than pixel-level identity, Kling produces the better scene.

    Load your own reference images into both from the AI video generator and run the same scene twice — identity judgments only count on your subject. Free credits daily.