Speech, voice and audio

    Speech-to-text

    Also called Transcription, ASR, Automatic speech recognition.

    Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.

    It runs at the front of more pipelines than people realise. Captions need it, dubbing needs it before anything can be translated, repurposing a long video into clips needs it to find where a topic starts. Almost any tool that appears to understand what a video is about is reading a transcript.

    Accuracy is not uniform across a file. Clear, close-miked speech in a common accent transcribes near-perfectly; overlapping speakers, background music, domain jargon and proper nouns are where errors cluster. Those errors then propagate, which is why a mistranslated brand name in a dub usually started as a mis-transcribed one.

    Word-level timing is the underrated output. A transcript that knows when each word was spoken is what makes karaoke-style captions and precise clipping possible; a transcript that is only text can be read but not aligned.

    In practice

    • Errors cluster on names, jargon and overlapping speech — proofread those, not the whole thing.
    • Timings matter as much as words if captions or clipping come next.
    • Isolating the vocal from music before transcribing measurably improves accuracy.

    The mistake to avoid

    Publishing auto-captions unread. The one word it got wrong is usually your product name, and it is on screen for the whole video.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.