Chatterbox TTS: a voice file, not a talking generate (5s, 10s, 30s, 3cr)
Chatterbox TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Chatterbox TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Chatterbox TTS is text-to-speech: natural, expressive narration with inline vocal tags like <laugh> and <sigh> woven into the script. The catalog calls it a solid default for UGC-style voiceovers. Category is text-to-audio. Audio is on. No image is required. Durations listed: 5s, 10s, 30s, 60s, 120s. The catalog lists 3 credits. There is no resolution because there is no picture.
Three credits, tags in the script, a two-minute ceiling. That is a VO file. It is not a creator's face moving.
Tags in the script, not on the face
<laugh> and <sigh> are performance marks inside the text. They change the file you hear. They do not open a jaw. If the cut is a talking UGC take, Chatterbox can still write the read — then a lipsync row has to wear it. If you skip that row, you have a laugh in the speakers and a still mouth in the frame. That is the UGC failure mode this model is named after.
Write the tags where a performer would breathe. Do not tag a visual joke you expected the model to act. There is no body.
The duration ladder is 5s, 10s, 30s, 60s, 120s. That is not Grok TTS (no durations, 15,000-character cap, 4 credits, 5 voices / 21 languages). Chatterbox is timed. A 5-second hook is a 5-second script on this row. A 2-minute story is 120s, not "paste until it stops." If you need 15 seconds, you do not have a 15s stop — you pick 10s and cut, or 30s and live with air. Match the read to the ladder before you generate.
The text-to-speech tool is the launcher. Add voiceover is how the file meets a plate. Best text-to-speech is the ranking. Chatterbox has no provider hub on this site; the model page is the row.
Two minutes is the long UGC read
120s is the ceiling. 60s is a minute of UGC VO. 30s is a typical ad read. 10s and 5s are hooks. Use the short stops to test the tagged performance cheaply — 3 credits listed is not an argument to start at 120s "because UGC is long." A bad sigh at 5s is information. A bad sigh at 120s is a minute you will not relisten to.
No image input. requires_image is false. Uploading a face here does nothing useful. Save the face for lipsync. Save the product still for image-to-video. Chatterbox only wants the script.
UGC-style in the description is a taste, not a platform export. You still caption. You still pick 9:16 in the picture pass. This row will not letterbox a voice file.
What 3 credits will not buy
A talking generate. Content type is audio.
A localisation matrix of 21 languages — that is a different TTS row in this catalog. Do not brief Chatterbox as a 21-language stack unless you are running separate jobs and checking the actual voice list in the product. Quote what the catalog says: tags, UGC-style narration, 5 / 10 / 30 / 60 / 120 seconds, 3 credits.
A substitute for lipsync when we see the mouth. Cheap VO makes the trap easier, not harder. The cheap version of a talking head is still two rows: this file, then a mouth pass.
What is the VO sitting on?
Look at the frame. Is there a mouth that should form these words?
If no — pack shot, screen, B-roll, faceless UGC — Chatterbox TTS at 5s or 10s is the right 3-credit file. Tag the laugh. Stay inside 120s.
If yes — creator talking to camera — generate the file if you want this voice, then lipsync. Do not lay 30s of tagged speech on a closed mouth and call it a talking ad.
FAQ
Is Chatterbox TTS a talking-head model?
No. Text-to-audio, 3 credits, durations 5s / 10s / 30s / 60s / 120s, inline tags such as <laugh> and <sigh>. You get a voice file. A mouth on screen is a later row.
How is this different from a TTS row with no duration list?
This slug publishes a duration ladder. Length is a stop you pick, not only a script you hope is short. 120s is the ceiling. Other TTS rows govern by character cap instead. Do not paste a two-minute script into a 5s stop and blame the tags.
Can I put <laugh> in the prompt of a video model instead?
Not as a substitute for this row. The tags are documented on Chatterbox TTS. A video model may ignore them or invent a visual laugh you cannot use. Write the performance here; move the mouth elsewhere if we see it.
Does it need a face upload?
No. Nothing visual is in scope. If you have a face that must speak, that upload belongs on image-to-lipsync, with this 3-credit file as the track — or on a talking generate that is not Chatterbox.