Guides

    Language tutors: Happy Horse is the talking row, captions are the muted row

    Do both. Native audio vs burned-in type are different jobs.

    Versely Team4 min read

    Do both. Native audio vs burned-in type are different jobs. Language-tutor video that only talks loses the mute scroll. Video that only captions never teaches a mouth. Happy Horse is the talking row. Captions are the muted row. They are not substitutes.

    A lesson clip has two audiences in the same feed: sound on, who need the mouth to match the language you named, and sound off, who need the words on screen in the first second. One generate does not finish both.

    Happy Horse is the mouth

    A talking shot is a clip whose job is that a visible mouth delivers a line, in a language you named, on a face you intend to keep. That is the tutor job: medium close-up, this face, this line, this language.

    Happy Horse 1.1 for multilingual talking shots is the router. Happy Horse 1.1 image-to-video animates a first-frame still into 1080p with native audio and multilingual lip-sync, 3 to 15 seconds; aspect is inferred from the image. The text-to-video sibling does the same from a prompt. The mouth is the feature. If you cannot see the mouth, you bought the wrong model.

    Same still, different scripts, each output lip-syncs to its language. One portrait becomes a per-market drill without five shoots and without a dub-then-repair pass. Write the language and the line into the prompt. Short sentences, one or two per clip. A 15-second monologue is three clips, not one paragraph.

    That is not a landscape with a voiceover, not an existing English take you must preserve (that is lipsync on footage you already have), and not a cloned founder timbre. Native multilingual delivery is fluent per language, not a clone of one voice.

    Native audio means picture and soundtrack from the same pass. Do not generate silent and slap TTS on a talking tutor unless the voice is a contractual object you cannot put in the prompt.

    Captions are the muted file

    Speech-to-captions is add captions to a video: Versely transcribes the spoken audio and burns styled, timed text. 21 BASIC presets at the standard credit rate, 9 DYNAMIC presets at 2×. 165 language codes. It transcribes what was spoken. It does not translate.

    That is the muted row. If the words are not on screen, the hook is gone. Captions are not a talking model. They run after the mouth exists.

    Set the caption language to the language actually spoken — a Spanish take captioned as en-US is a bad transcript, not a translation. Accuracy follows audio clarity. Preview a DYNAMIC look on the first seconds before a full render.

    A claim nobody says — "A2, 12 minutes, next Tuesday" — is not transcription. Write it. Burn a fixed line with text overlay, or several timed lines with timed text overlays. Do not ask Happy Horse to typeset the syllabus onto a moving face.

    Do both; never as substitutes

    The talking row teaches pronunciation. The muted row keeps the file alive when the phone is silent. Shipping only Happy Horse is how a good mouth dies in mute autoplay. Shipping only captions is type with no mouth to learn from.

    Lock the still, face large. Generate the talking take on Happy Horse — name the language, quote the line. Caption that file in the spoken language. Overlay schedule and level as typed text, not as generated type in the mouth pass.

    If you need the same voice next week, that is TTS on a lipsync pipeline. If you need the same face in three languages, that is Happy Horse three times from one still. Do not mix both on one mouth.

    FAQ

    Can captions replace the talking generate for language learners?

    No. Captions are burned-in type from a transcript. They do not move a mouth. Learners who have sound on still need the talking row. Do both.

    Does Happy Horse translate my English script?

    No. Write the line in the target language. The model will lip-sync a bad translation with perfect confidence. Native speakers still check the script.

    Will add-captions-to-video translate Spanish speech into English type?

    No. It transcribes. The caption language should match the language spoken. Translation is a different editing task.

    Why not generate silent and caption a TTS bed?

    On a visible mouth, silent-plus-TTS throws away the only free lipsync you were going to get. Use that path when the voice is a locked brand object and the picture must be driven from audio. A tutor drill is usually the other job: native talking take, then captions.