Guides

    Inworld TTS: a voice file, not a talking generate (3cr)

    Inworld TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team3 min read

    Inworld TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Inworld TTS is Inworld's text-to-audio row: 3 credits, audio out, no picture. The catalog job is "High-quality text-to-speech powered by Inworld AI. Supports a curated library of expressive voices with multilingual capability." Features listed: Text-to-Speech, Expressive, Multilingual, Voice Library. You pick a voice from that library. You do not upload a clone sample on this row, and you do not get a mouth.

    A library voice, 3 credits

    This is not a clone generate. It is a curated library. Pick the expressive voice, write the line, take the file. The text-to-speech tool is the door. Text-to-speech is the mode. Three credits is the catalog price. There is no resolution and no duration ladder because nothing is on screen.

    Multilingual capability is listed. That is a VO job: same script, another language, same library pick (or a better match in that language). It is not a dubbed talking-head pipeline by itself. The mouth still needs a lipsync row if we can see it.

    Closed-mouth clips stay closed

    If the cut is a face in frame, this file is incomplete. AI lipsync or a talking generate is the rest. If the cut is b-roll, this file is complete. Add a voiceover and stop. Laying Inworld under a silent jaw is how a cheap VO becomes an expensive reshoot of the mouth.

    Do not start here when the picture does not exist. Start with the plate. Come back for the library voice when you know whether the jaw is in shot. The Inworld provider roster is the rest of that brand. This page is only the TTS library row.

    Write the line, pick the voice

    A useful Inworld pass is a script, a library choice, and a decision about the picture. Expressive is a listed feature — use it for a line that has to land, not for a paragraph the model will flatten. If you need a specific person's voice, that is a clone row, not this library. If you need a designed voice from a description, that is a voice-design row.

    Three credits is not a talking clip. It is a voice file. Keep the categories honest and the mouth will not surprise you at upload.

    FAQ

    Does Inworld TTS make a talking video?

    No. Content type is audio. Category is text-to-audio. You get a voice file from a curated library.

    Is this a voice clone?

    No. This record is a library of expressive voices, not an upload-a-sample clone. Clone is a different Inworld row if you need a specific person.

    Do I need an image?

    No. This row does not require an image. An image would not open a mouth. Use lipsync when the face is in frame.

    What are the 3 credits for?

    The catalog price for this text-to-speech generate. Not a 4K avatar, not a scene, not a duration.