Spoken clip in, transcript out: transcribe_audio does not burn captions
Transcribe audio to text is one agent job. transcribe_audio returns words. Burning styled captions on a video is a different job.
Spoken clip in. Written transcript out. Transcribe audio to text is transcribe_audio (Cartesia ink-whisper) on an audio file. You get text. You do not get a styled, timed caption burn on a video. Using this job as a captioner, or captioning a video when you only needed a memo in a doc, is how you pay a finishing pass for a string, or sit on a wall of text when you needed burned-in type.
The tool requires audio_url. Optional language. Lightweight processing — check your plan — not a video-model generation. Get a video transcript is the neighbouring editing task when the source is a video you intend to stay a video.
Words are the product
"Transcribe this voice memo." "Give me a written transcript of this recording." That is the list. If the ask is "add glass captions to this Reel," the job is transcribe and caption my video — speech plus VEED presets burned onto picture. Transcript first, burn later, is correct when you need to edit the words. Burn first when you never needed a doc.
Do not send transcribe_audio a video URL and expect a captioned MP4. Do not send it a song and expect lyrics aligned to bars; lyrics generation and stem splits are other tools. Do not use a full video breakdown (analyze_video) when you only wanted the speech as text. Analysis extracts frames, beats, on-screen type, and an optional transcript. That is a heavier pass.
If you needed a script to re-voice, transcribe, edit the text, then generate speech. Transcribe is the recovery of words that already exist. It is not TTS.
Language is a parameter
If you know the spoken language, say it. Regenerating the transcript because you forgot language is a second process for a field the first call accepted. The output is still text. Captions still need the caption job.
FAQ
Will transcribe_audio burn subtitles onto my video?
No. Text out. Caption the video is the burn, with a preset. Paying for captions when you needed a pasteable transcript is the wrong finishing spend.
Can I transcribe a video file with this job?
This tool wants audio. If you have a talking video and you need on-picture captions, use the caption capability. If you only need the words, pull audio or use the transcript editing task.
Is this the same as analyze_video with include_transcript?
No. analyze_video is a breakdown: frames, beats, on-screen text, optional speech. Use it when you need the whole description. Use transcribe_audio when the only artifact is the words.
Does transcription generate a new voiceover?
No. It writes what was already spoken. A new read is generate_speech.