transcribe_audio runs Cartesia's ink-whisper speech-to-text model on an audio clip and returns a written transcript. It's useful any time you need the words in text form — repurposing a video's dialogue as a blog post, pulling quotes, or drafting a script for a translated version.
This is different from add-subtitles-automatically, which burns the transcript directly onto the video as timed captions. transcribe_audio gives you plain text you can edit, copy, or feed into another tool first.
Powered by
Transcribe an audio clip to text (Cartesia ink-whisper speech-to-text). Use when the user wants a spoken clip turned into a written transcript.
What to tell the agent
Versely's agent maps this plain-English request directly onto transcribe_audio. You don't need to know the parameter names — just describe what you want.
“Transcribe the audio from this video to text.”
How it works
1. Point to the audio
Supply an audio_url — pulled from a video's audio track or a standalone recording.
2. Set the language
Pass language if you want to specify it explicitly rather than relying on auto-detection.
3. Run the transcription
Cartesia's ink-whisper model converts the speech to text.
4. Reuse the text
Edit it, translate it, turn it into a script for dub_video, or feed it back in as the source for add_timestamped_captions.
What it costs
Billed per audio clip transcribed, based on length. Ask the agent for estimate_cost before transcribing a long file.
The formula behind that number: What does a finished 30-second AI ad cost end to end? — sums four unrelated meters, so no single rate predicts the total.
Limits & things to know
- Accuracy depends on audio clarity — background noise, overlapping speakers and heavy accents reduce accuracy.
- This returns plain text, not a burned-in caption track — pair it with a captioning tool if you want the words on screen.
Who uses this
- Repurposing video dialogue as blog or social copy
- Pulling exact quotes from an interview
- Drafting scripts before a translation or dub pass
- Searchable archives of spoken content
Frequently asked questions
Does this add subtitles to my video?+
No — transcribe_audio returns plain text only. To burn transcribed subtitles onto the video itself, use add-subtitles-automatically instead.
What speech-to-text engine does it use?+
Cartesia's ink-whisper model.
Can I specify the spoken language?+
Yes — pass a language code, or leave it to auto-detection.
Related Versely tools
Related editing jobs
Get a Video Transcript inside Versely
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.