Guides

    Inworld TTS 2: a voice file, not a talking generate (5cr)

    Inworld TTS 2 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team4 min read

    Inworld TTS 2 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Inworld TTS 2 is Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language style steering for fine-grained delivery control. Category: text-to-audio only. Content type is audio. It does not require an image. No duration chips. No resolution. Display price: 5 credits. Featured. This is a research-preview speech engine, not a talking-head generate.

    One hundred languages, still a wav

    100+ languages is the headline. Style steering is the other: you describe the delivery in natural language and the model performs it. Expressive, fine-grained, still a file. A French take and a Korean take are two wavs. They are not two faces. They are not one lipsync pass that "just switches language."

    The catalog does not list voice-clone on this slug. Cloning is a different Inworld row if you need it. TTS 2 is text-to-audio: script in, speech out, 5 credits. Compare that to a clone-capable speech row only when you actually need a clone. Do not assume preview power includes a likeness.

    The door is AI text-to-speech. The Inworld roster is TTS generations plus cloning on other slugs. Inworld TTS is the earlier speech page if you are not on the research preview. Language coverage across engines is on voice-over. Ranked speech: best text-to-speech model.

    Style steering is delivery, not picture

    Natural-language style steering is easy to over-read. "Tired, close to the mic, almost a whisper" is a performance note. It is not a lighting setup. It does not turn a product still into a presenter. Audio tags and style notes punch the take. They do not track lips.

    Research Preview is in the description. Treat the row as a powerful expressive engine with that label attached, not as a finished talking-avatar product. 5 credits is the display price (per 1,000 characters on the matrix). Plan the speech job at 5 credits. Plan the mouth separately.

    If the paragraph is a wall, TTS 2 will flatten the same as any other engine. Chunk it. That is narration going flat, not a reason to jump to a video model.

    Closed mouth means you picked wrong

    The visual test is the whole page. If we see teeth, this 5-credit file is not the last generate. Pick a talking or lipsync row that takes a plate. AI lipsync is that door. Generate the line here — 100+ languages is why you might — then drive the face on a lipsync model.

    If we do not see the mouth, you are done when the wav lands. Localization is a speech problem until the cut is a talking head. Then it is speech plus mouth, twice the rows, two credit numbers.

    Five credits like other speech, different engine

    5 credits is the same display figure as some other TTS rows. The spec is not the same: 100+ languages, style steering, research preview, text-to-audio only, Inworld. That is enough to pick it. It is not enough to skip the plate.

    Do not lay this on a closed mouth and call it a presenter. That is a routing miss with a multilingual wav underneath.

    FAQ

    Does Inworld TTS 2 output a talking video?

    No. Text-to-audio, 5 credits, no image, no resolution. You get a voice file. The mouth is a lipsync or talking-head row.

    Is voice cloning included on TTS 2?

    Not on this slug. Categories are text-to-audio only. If you need a clone, that is another Inworld row. TTS 2 is script-to-speech with style steering and 100+ languages.

    What does "Research Preview" mean for production?

    It is in the catalog description. Use it as Inworld's most powerful expressive engine with that label. It still writes audio. It still does not invent a face.

    When do I leave TTS 2 for lipsync?

    When the cut shows a mouth. Generate the line here, lock the still, then run a lipsync model. If the cut is VO over picture, 5 credits and stop.