Speech, voice and audio

    Forced alignment

    Also called Word-level timing, Auto-timing.

    Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.

    It is a different problem from transcription, and easier. Transcription has to work out what was said; alignment already has the text and only has to place it on the timeline. Because the answer is constrained, alignment is accurate to a fraction of a second even where recognition would struggle.

    This is the machinery behind captions that highlight word by word as they are spoken. Sentence-level subtitles need only rough timings; a caption that pops each word on the beat needs to know where every word starts and ends, and that is alignment output.

    It also enables precise cutting. Knowing the timestamp of a specific word means you can trim a sentence exactly, remove a filler word without an audible seam, or split a long recording at natural boundaries rather than at arbitrary intervals.

    In practice

    • Feed it a corrected transcript — alignment inherits your typos as spoken words.
    • Word-level timing is what animated caption styles consume.
    • Music and overlapping speech degrade timing accuracy the same way they degrade transcription.

    The mistake to avoid

    Assuming captions are late because the alignment is wrong. More often the caption style has a display duration that outlives the word it belongs to.

    Go deeper

    Timed Captions From Speech: How Auto-Timing Works

    How AI turns speech into perfectly timed captions: word-level timestamps, chunking logic, failure modes, and when to override auto-timing manually.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.