Guides

    Spoken clip in, transcript out: transcribe_audio does not burn captions

    Transcribe audio to text is one agent job. transcribe_audio returns words. Burning styled captions on a video is a different job.

    Versely Team3 min read

    Spoken clip in. Written transcript out. Transcribe audio to text is transcribe_audio (Cartesia ink-whisper) on an audio file. You get text. You do not get a styled, timed caption burn on a video. Using this job as a captioner, or captioning a video when you only needed a memo in a doc, is how you pay a finishing pass for a string, or sit on a wall of text when you needed burned-in type.

    The tool requires audio_url. Optional language. Lightweight processing — check your plan — not a video-model generation. Get a video transcript is the neighbouring editing task when the source is a video you intend to stay a video.

    Words are the product

    "Transcribe this voice memo." "Give me a written transcript of this recording." That is the list. If the ask is "add glass captions to this Reel," the job is transcribe and caption my video — speech plus VEED presets burned onto picture. Transcript first, burn later, is correct when you need to edit the words. Burn first when you never needed a doc.

    Do not send transcribe_audio a video URL and expect a captioned MP4. Do not send it a song and expect lyrics aligned to bars; lyrics generation and stem splits are other tools. Do not use a full video breakdown (analyze_video) when you only wanted the speech as text. Analysis extracts frames, beats, on-screen type, and an optional transcript. That is a heavier pass.

    If you needed a script to re-voice, transcribe, edit the text, then generate speech. Transcribe is the recovery of words that already exist. It is not TTS.

    Language is a parameter

    If you know the spoken language, say it. Regenerating the transcript because you forgot language is a second process for a field the first call accepted. The output is still text. Captions still need the caption job.

    FAQ

    Will transcribe_audio burn subtitles onto my video?

    No. Text out. Caption the video is the burn, with a preset. Paying for captions when you needed a pasteable transcript is the wrong finishing spend.

    Can I transcribe a video file with this job?

    This tool wants audio. If you have a talking video and you need on-picture captions, use the caption capability. If you only need the words, pull audio or use the transcript editing task.

    Is this the same as analyze_video with include_transcript?

    No. analyze_video is a breakdown: frames, beats, on-screen text, optional speech. Use it when you need the whole description. Use transcribe_audio when the only artifact is the words.

    Does transcription generate a new voiceover?

    No. It writes what was already spoken. A new read is generate_speech.