Lay B-roll under an AI voice track
Record TTS, generate product I2V, lay it under the voice, then caption. A talking head is optional.
Every guide, comparison and workflow we’ve published on Voiceover.
35 articles — page 1 of 2
Record TTS, generate product I2V, lay it under the voice, then caption. A talking head is optional.
ElevenLabs, Murf, OpenAI TTS, Versely, and Cartesia all voice a 15-second ad. Pick by clone, timeline, API cents, credits, or Sonic punch.
ElevenLabs, Murf, Hume, WellSaid, Versely, and Descript cover weekly VO. Pick by clone, studio, emotion, minutes, or a credit TTS row.
Cartesia, Inworld, Murf, Versely, WellSaid, and OpenAI TTS all read a script. Pick by latency, style, studio timeline, credits, actor voices, or API cents.
Descript, Adobe Enhance, CapCut, ElevenLabs, Versely, Murf, and WellSaid cover voiceover with no booth. Clean a take or generate one.
Faceless UGC is B-roll under a voice. A generated face is still a talking head, not a faceless cut.
Music that fights the voice kills retention. Duck hard under narration and stop romanticizing a hot mix.
TTS is ceil(chars ÷ 1000 × 12) credits; pacing, pauses and emotion change runtime without changing the bill.
Add music or a voiceover to a video is one agent job. attach_audio_to_video lays an existing file. Generating speech or music is a different job.
ElevenLabs Multilingual writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Grok TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
MiniMax Speech writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
A full mix under VO is a fight. Get stems, or generate the bed without a lead, then duck.
/free-tools/script-timing-estimator never spends a credit. Do not pay a model to do an in-browser file job.
Write and generate a voiceover is one agent job. generate_speech returns a spoken file. Laying it on picture, or casting two speakers, is a different job.
generate_speech plus attach_audio_to_video lays a script on a clip. Narrating a take you will regen, or narrating before picture lock, pays for a read you will orphan.
AD that ducks the whole mix is unusable; AD under the music is inaudible. Match description to dialogue, keep HI and VI mixes apart, and check on a phone speaker.
TTS and clones hiss in a way live VO does not. A split-band de-esser setup, and when the real fix is a different take or engine rather than more processing.
Pass a speech-to-noise check and a mono phone test before sign-off, and fix masked words with EQ and reverb moves rather than raising the VO.
Resample every 44.1 kHz generate on ingest, use a real anti-alias filter, and keep the project at 48 kHz / 24-bit so export does not click.
Fold 5.1 and Atmos beds with centre-channel priority, discard the LFE, and check the result on a phone speaker so the words still survive.
Present tense, character IDs, and dialogue-gap timing for audio description, plus a timed script template you can fill for a 30-second spot.
Prosody decays as the model loses sentence structure. A chunking strategy, punctuation that restores emphasis, and joins you cannot hear.
A real or licensed cameo plus an original voice track is the cheapest move from template-shaped to authored. The layering order and how to keep it consistent.