Guides

    Transcribe the keeper, not the miss you are about to replace

    transcribe_audio (Cartesia ink-whisper) returns plain text. Running it on a throwaway soundtrack, or before picture lock, pays for a script of words you will never reuse.

    Versely Team3 min read

    A transcript is the words, as text. transcribe_audio runs Cartesia's ink-whisper model on a clip and gives you a document — for a blog, a quote card, a corrected caption pass, a translation draft. It is not burned-in subtitles. It is also not free to spray across every sample in the bin. A miss has the wrong words. A recut drops lines. You paid for a script of a file you will delete.

    Get a video transcript is the job. Billed per clip, by length. estimate_cost before a long file. Accuracy follows clarity: noise, overlap, heavy accents. That is a reason to isolate a keeper, not a reason to transcribe a maybe "to see what it said."

    Text is not a caption track

    If you needed type on screen, automatic subtitles burn timed cues. Transcribe when you need to edit the words first: names, SKUs, legal. The honest pipeline is: lock soundtrack → transcribe → correct → burn. Transcribing a scratch VO so you can "start the blog" produces a draft you will not trust, and a second transcript later.

    People transcribe first because text feels like progress. In a generated pipeline the progress is the take. If the talking head is getting regenerated, the transcript is fan fiction about a performance that will not ship.

    This is also the wrong first step for translation. Translate a signed English (or source) document, or dub a signed file. Do not transcribe take four in a stack of six.

    When the pass belongs

    • The audio is the mix you will ship.
    • You need quotes, a blog, a QA script, or a caption file you will correct.
    • You are about to subtitle a long keeper and names have to be exact.

    Then you have a document worth keeping. Pair it with a caption tool if the destination is a player. Do not confuse the two bills.

    The voice cloning surface sits nearby because a clean transcript is sometimes how you check what a clone will be asked to say. Still later. Still on keepers.

    A cheap alternative to a premature transcript

    If you only needed to know whether the line landed, watch the take. Your ears are free. Pay ink-whisper when someone else has to read the words without the video.

    FAQ

    Can I reuse a transcript on the regen if the script was the same?

    Only if the regen actually said the same words at the same times. Generated takes ad-lib. Even a "same prompt" sample is a new performance. Transcribe the file you will publish.

    Should I transcribe before isolating noisy audio?

    If the keeper is noisy and you need accuracy, isolate first. Transcribing the bed is how you get lyrics mixed into the VO.

    Is this how I get an SRT?

    It returns plain text, not a timed sidecar. Timed burn-in is the caption tools. If you need a sidecar master, plan for that format — do not assume transcribe_audio is it.

    Why not transcribe every generation as logging?

    Because it is billed, and the log of a deleted take is not an archive you wanted. Log prompts and shot ids. Transcribe keepers.