Add multi-speaker dialogue as audio
Generate one conversation with a voice per alias, then attach it. Two solo reads stitched still sound like two monologues.
Generate one conversation with a voice per alias, then attach it. Two solo reads stitched still sound like two monologues.
A silent row returns picture. Write the script, generate speech, replace the empty track. Do not prompt dialogue on Turbo.
Hume Octave 2 is not on Versely, and Hume's API closes 13 November 2026. Cartesia Sonic 3.5 is 4 credits per 1,000 characters here.
Generate silent, speak with Gemini 3.1 Flash TTS (4cr), then replace the plate's audio. Do not mix.
Kling Turbo (12cr, silent), then Gemini 3.1 Flash TTS (4cr), then burn-in. Native audio is wasted when the line must be exact.
NY §396-b skips audio-only ads and translation-only AI. A talking-head video of a fake person is not exempt. TTS and video split.
Cartesia Sonic and Hume Octave 2 both read scripts. Octave leans emotion controls. Sonic leans low-latency clarity. Pick by the performance.
Voice cloning without rights is a lawsuit shaped like a convenience. Own the consent.
Gemini 3.1 Flash TTS supports inline audio tags for sighs and laughs. Use them sparingly or the read feels cartoon.
Inworld TTS 2 and Gemini 3.1 Flash TTS are both voice files. Style steering and voice counts differ. Neither is a talking generate.
Wan 2.7 T2V and I2V rows in catalog carry null audio. Plan TTS and lipsync as separate steps when speech is required.
If we see a mouth, it is lipsync or Veo. TTS is audio-only.
Firefly Speech Model is the licensed path. ElevenLabs is offered as an option with different terms. Read which one ran.
Inworld TTS 2 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Cloning a noisy take copies the noise. Isolation is the first step, not a polish.
If the model can speak in-shot, do not generate silent and slap TTS on unless you need a locked brand voice.
TTS style lock is the thing that makes 30 episodes sound like one show. Set it once.
Write and generate a voiceover is one agent job. generate_speech returns a spoken file. Laying it on picture, or casting two speakers, is a different job.
generate_speech plus attach_audio_to_video lays a script on a clip. Narrating a take you will regen, or narrating before picture lock, pays for a read you will orphan.