Recording voiceover yourself — or hiring someone to — is the slowest step in most edits. generate_speech turns a script into spoken audio, with emotion and style controls so it doesn't read flat, then attach_audio_to_video lays that narration onto your footage.
Decide up front whether the voiceover should replace the video's existing audio entirely (mode 'replace') or sit alongside it (mode 'mix', e.g. narration over ambient B-roll sound).
Powered by
Generate speech audio from text (Text-to-Speech). Use when the user needs a voiceover, narration, or spoken audio for their content. Perfect for UGC workflows, ad scripts, and podcast clips. Use `emotion` and/or `style_instructions` to make the delivery expressive and match the content's mood instead of defaulting to a flat read — pick whichever the target model supports.
What to tell the agent
Versely's agent maps this plain-English request directly onto generate_speech. You don't need to know the parameter names — just describe what you want.
“Generate a warm, conversational voiceover reading this script, then attach it to the video, replacing the original audio.”
How it works
1. Write the script
Paste the narration text. Add emotion or style_instructions if you want the delivery to match the content's mood rather than a flat read.
2. Pick a voice and model
generate_speech takes a model and voice_id — choose a stock voice, or a cloned one if you've already run clone-your-voice-for-videos.
3. Attach it to the video
Use attach_audio_to_video with mode 'replace' to make the narration the only audio track, or 'mix' to layer it over existing ambient sound.
4. Render
You get back a finished MP4 with the narration synced to the video's timeline from the start.
What it costs
generate_speech is billed per generation depending on model and script length; attach_audio_to_video adds a flat fee to lay the track onto the video. Ask the agent for estimate_cost before generating a long script.
The formula behind that number: What does an AI voiceover cost for my script? — meters the characters you type, not the seconds you get back.
Limits & things to know
- generate_speech produces a separate audio file — it isn't automatically timed to specific on-screen actions beyond starting at the video's beginning.
- For dialogue between multiple distinct voices, use add-multi-speaker-dialogue-to-video instead of chaining several single-voice calls.
- Emotion and style_instructions support varies by which TTS model you pick.
Who uses this
- Explainer and tutorial videos
- Faceless YouTube channels
- Product demo narration
- Documentary-style B-roll
Frequently asked questions
Can I control how expressive the narration sounds?+
Yes — generate_speech accepts emotion and style_instructions so the delivery can match your content's mood instead of defaulting to a flat read.
Should the voiceover replace or sit alongside my video's audio?+
Use attach_audio_to_video's mode 'replace' if the voiceover should be the only track, or mode 'mix' to keep the video's existing sound (e.g. ambient B-roll noise) playing underneath it.
Can I use my own cloned voice for the narration?+
Yes — clone a voice first with clone_voice_from_audio, then pass its voice_id into generate_speech.
What if I need two different voices talking to each other?+
For scripted dialogue between multiple distinct voices, use generate_multi_speaker_speech instead — see add-multi-speaker-dialogue-to-video.
See it in a workflow
Related Versely tools
Related editing jobs
Add a Voiceover to a Video inside Versely
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.