Guides

    Audiobook ads: TTS is not the talking shot of the author

    If we see a mouth, it is lipsync or Veo. TTS is audio-only.

    Versely Team4 min read

    If we see a mouth, it is lipsync or Veo. TTS is audio-only.

    An audiobook ad that needs the author on camera is a talking shot. An audiobook ad that needs the chapter in someone's ear is a voice file. Mixing the two — a still of the author, a TTS bed, and a mouth that is chewing while the sentence has already moved on — is how a launch teaser feels like a cousin of the narrator. Native audio versus TTS on the same brief is the routing rule. This page is the author-teaser version: do not pretend a speech model is a face.

    TTS has no mouth

    AI text to speech reads a script. You paste the line, you pick or describe a voice, you get audio. The meter is per 1,000 characters of the text you send, including spaces and punctuation. Playback speed does not change the bill. Cutting words does.

    That file does not know about a face. It does not move lips. If you lay it under a silent generate of the author, the mouth will not match unless you add a lipsync pass — a third product. If you lay it under a native-audio generate, you now have two voices fighting one face.

    Use TTS when the shot has no mouth: cover still, hands on the page, a product turn of the hardback. Audition one sentence across engines before you commit the teaser. Style-lock the voice if episode 12 has to sound like episode 1. Native audio will give you a plausible speaker every time, and a different plausible speaker every time. That is a recast for a series.

    If we see lips, pick the picture job

    Native audio is a video-model feature. Veo 3.1 and the rest of the native-audio shortlist generate speech with the face that is saying it. Write the line in the prompt, in quotes. Leaving sound unspecified is how you get a saxophone on a literary ad.

    Lipsync is how you keep a contractual voice on a contractual face. AI lipsync takes a still (or a clip) and a voice file — TTS, a clone, a studio take — and drives the mouth from the audio. That is the honest talking author when the voice is the named author or a licensed narrator, and the video model cannot take that voice as an input.

    The jacket still is not a talking shot. Animate it if you want motion without speech. Put TTS under it if the author is not speaking on camera. Do not ask the still to lip-flap from a speech model you never connected.

    Native audio is picture and soundtrack from the same pass. TTS is a voice you control. Pick one per mouth.

    One source per mouth on a teaser

    The failure that looks finished in the timeline: a Veo take of "the author," the native track ducked, the TTS chapter sample dropped on top. The lips still belong to the first voice. Listeners may not name the artefact. They will feel it.

    Do not regenerate native until the voice "sounds like" the author — the model does not own that timbre. Do not generate silent "so we can add TTS later" on a talking close-up; you threw away the only free lipsync you were going to get. Native on the A-roll and TTS on the same person's B-roll is two people.

    If the teaser is a chapter in someone's ear over the cover, TTS is the job. If the teaser is the author saying one line to camera, native audio or lipsync is the job. Caption after the voice is decided. Captions are not a third audio path.

    FAQ

    Can I put a TTS chapter sample under a still of the author?

    Yes, if the mouth is not the shot. A held jacket photo plus a locked narrator is TTS's home turf. The moment the mouth moves, you are in lipsync or native audio.

    Does native audio mean I should never use TTS on an author ad?

    No. Off-screen narration over hands, a cover, or b-roll is TTS. The rule is one source per mouth, not one source per video.

    When is silent-plus-TTS the right talking shot?

    When the voice is a contractual object — the named author, a licensed narrator — and the video model cannot take that voice as an input. Then drive the mouth from the audio with lipsync. That is not a "we'll fix it in captions" job.

    Is TTS billed by how long the teaser runs?

    No. Speech models bill by the characters you send. A 900-word sample costs the same whether the read is brisk or slow. Cut words to cut the bill. Picture is a different meter.