Guides

    ElevenLabs Multilingual: a voice file, not a talking generate (6cr)

    ElevenLabs Multilingual writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team4 min read

    ElevenLabs Multilingual writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    The catalog card: multilingual text-to-speech powered by ElevenLabs Multilingual v2. Content type audio. Category text-to-audio. 6 credits. Catalog audio is true — this is the audio. No resolution. No duration enum. No image required. You get a voice file. You do not get a face.

    Laying that file on a generate whose lips never opened is how you ship a talking ad with a closed mouth. Laying it on a native-audio clip is how you ship two voices on one face.

    The output is a track

    ElevenLabs Multilingual is a text-to-speech row. You write the script. You pick the voice. You get speech. Delivery tags belong in the script: [excited], [whispers], [laughs], [sighs], [sarcastic], [curious] — short, sparse, immediately before the phrase they color. Tagging every clause is how the read becomes a cartoon.

    This is not Veo. This is not Fabric. This is not O3 Pro I2V with generate-audio on. Those rows make picture (and sometimes a mouth). This row makes a wav you can lay under b-roll, a product turn, or a landscape. The AI text to speech tool is the door. The best text-to-speech model ranking is the shelf. Add voiceover to video is the edit job after you have the file.

    Multilingual v2 is the point of the name: the same engine across languages, not a talking avatar that happens to be bilingual. If the brief is "the presenter speaks Spanish on camera," you still need a mouth pipeline. This card will happily read Spanish into a file while the picture stays mute.

    Closed mouth, open file

    Use this row when:

    • The face is not the shot, or there is no face.
    • The voice has to be the same person next week (a chosen ElevenLabs voice, not a random native-audio extra).
    • You need a read you can time, duck, and caption without regenerating picture.

    Do not use it when:

    • We see the mouth. That is AI lipsync or a talking generate with native audio.
    • You hoped the video model would stay silent and it did not. Fix the video row; do not stack this on top of a spoken extra.

    Six credits is the catalog figure for the card. Billing is also listed per 1,000 characters on the price matrix. Either way you are buying speech, not pixels. A long script is a long bill; it is still cheaper than regenerating a 70-credit clip to change one clause.

    Keep picture and voice on separate contracts

    If the picture already spoke, you do not "fix" it by dropping Multilingual v2 on the timeline. You pick a silent plate, or you convert an existing recording on ElevenLabs Voice Change, or you lipsync. The native-audio versus TTS rule is this brief. This page is only the TTS side: a voice file, 6 credits, no mouth.

    FAQ

    Can ElevenLabs Multilingual make a talking-head video?

    No. It writes audio. Pair it with b-roll, or send the track into a lipsync / talking row that takes audio or a script plus a plate. This card has no picture output.

    Why not generate a silent Sora clip and always lay this on top?

    Because Sora 2 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose. Use this TTS row when the picture is actually mute, or when the mouth is not in frame.

    Do I write emotion in a separate style field?

    Not on ElevenLabs. Delivery is in-text tags immediately before the phrase. Sparse beats a tag on every sentence.

    Is 6 credits a per-second video rate?

    No. This is an audio row. You are not buying 4s of 720p. You are buying a voice file from Multilingual v2.