Inworld TTS 1.5 Max: a voice file, not a talking generate (4cr)
Inworld TTS 1.5 Max writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Inworld TTS 1.5 Max writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Inworld TTS 1.5 Max is Inworld Realtime TTS 1.5 Max — the catalog's own description calls it the #1 ranked Inworld model, the balance of quality and speed, expressive, multilingual (15 languages), with voice library support. Category is text-to-audio. Four credits. Content type is audio. audio on the row is null because this is not a video flag; the output is the voice file. No resolution. No duration enum.
Four credits, fifteen languages, still not a face
Multilingual is the reason to be on this slug rather than a monolingual compact TTS. Fifteen languages, a voice library, expressive delivery. That is a narration and localisation file. It is not a presenter. It does not open a mouth.
If we can see lips, this file laid on a silent generate is a desync waiting for comments. Options that own the mouth: a native-audio video model with the line in the prompt, or a lipsync pass that takes this track as the driver. Off-screen VO over b-roll is the clean Inworld job.
The AI text-to-speech tool is the door. Inworld's provider hub is the family. Best text-to-speech model is the ranked list. This page is only 1.5 Max: 4 credits, 15 languages, voice library, audio file out.
Voice library is not lipsync
Library support means you pick a voice that survives a series. Same timbre next Tuesday. That is the TTS contract. Native video audio will give you a plausible speaker every time and a different plausible speaker every time. If the series bible names a voice, lock it here. Then decide whether the mouth is in frame.
Inline delivery on Inworld sits at the word or clause it modifies — mid-sentence, not only at the top of the line. Tag the shift where it happens. Do not wrap the entire paragraph in one mood and call it expressive. Expressive is sparse.
There is no clip length because length follows the script. Write the line you will use. A localisation pass is the same voice library, a new script, the same 4-credit row — not a new talking-head generate in each language.
Fifteen languages still need a mouth row
Do not:
- Generate Kling Turbo (audio off) and drop Inworld on a chewing face
- Generate Wan V2.6 (audio on) and duck the native stem for Inworld on the same mouth
- Ask Inworld to "sound like the avatar" and skip the lipsync row
Do:
- Inworld for off-screen narration and multilingual VO
- Add voiceover when the plate is not a talking shot
- Lipsync or a talking model when the plate is a talking shot
- Captions after the voice is decided, including for the 14 languages that are not the source
Native audio versus TTS is the brief-level rule. Inworld 1.5 Max is the TTS row with a 15-language library at 4 credits. Use it as a file. Do not use it as a person.
FAQ
Does Inworld TTS 1.5 Max generate video?
No. Content type is audio. Category is text-to-audio. Four credits buys a voice file. A talking presenter is a different category.
What does "15 languages" actually mean for production?
The catalog description: multilingual (15 languages) with voice library support. Lock the voice, then run the localised script on this same row. It does not localise a mouth. If the face is in frame, plan lipsync or a talking model per language.
Is this a clone of my founder?
Voice library support is not automatically a clone of a named person. If you need a contractual founder voice, that is a clone pipeline, then this kind of TTS or a lipsync driver. Do not treat "expressive" as "this specific human."
Can I put Inworld audio on any silent clip?
On clips where we do not see a speaking mouth, yes. On clips where we do, only through lipsync. A closed mouth plus a fluent 15-language read is the tell.