Guides

    Chatterbox TTS: a voice file, not a talking generate (5s, 10s, 30s, 3cr)

    Chatterbox TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team5 min read

    Chatterbox TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Chatterbox TTS is text-to-speech: natural, expressive narration with inline vocal tags like <laugh> and <sigh> woven into the script. The catalog calls it a solid default for UGC-style voiceovers. Category is text-to-audio. Audio is on. No image is required. Durations listed: 5s, 10s, 30s, 60s, 120s. The catalog lists 3 credits. There is no resolution because there is no picture.

    Three credits, tags in the script, a two-minute ceiling. That is a VO file. It is not a creator's face moving.

    Tags in the script, not on the face

    <laugh> and <sigh> are performance marks inside the text. They change the file you hear. They do not open a jaw. If the cut is a talking UGC take, Chatterbox can still write the read — then a lipsync row has to wear it. If you skip that row, you have a laugh in the speakers and a still mouth in the frame. That is the UGC failure mode this model is named after.

    Write the tags where a performer would breathe. Do not tag a visual joke you expected the model to act. There is no body.

    The duration ladder is 5s, 10s, 30s, 60s, 120s. That is not Grok TTS (no durations, 15,000-character cap, 4 credits, 5 voices / 21 languages). Chatterbox is timed. A 5-second hook is a 5-second script on this row. A 2-minute story is 120s, not "paste until it stops." If you need 15 seconds, you do not have a 15s stop — you pick 10s and cut, or 30s and live with air. Match the read to the ladder before you generate.

    The text-to-speech tool is the launcher. Add voiceover is how the file meets a plate. Best text-to-speech is the ranking. Chatterbox has no provider hub on this site; the model page is the row.

    Two minutes is the long UGC read

    120s is the ceiling. 60s is a minute of UGC VO. 30s is a typical ad read. 10s and 5s are hooks. Use the short stops to test the tagged performance cheaply — 3 credits listed is not an argument to start at 120s "because UGC is long." A bad sigh at 5s is information. A bad sigh at 120s is a minute you will not relisten to.

    No image input. requires_image is false. Uploading a face here does nothing useful. Save the face for lipsync. Save the product still for image-to-video. Chatterbox only wants the script.

    UGC-style in the description is a taste, not a platform export. You still caption. You still pick 9:16 in the picture pass. This row will not letterbox a voice file.

    What 3 credits will not buy

    A talking generate. Content type is audio.

    A localisation matrix of 21 languages — that is a different TTS row in this catalog. Do not brief Chatterbox as a 21-language stack unless you are running separate jobs and checking the actual voice list in the product. Quote what the catalog says: tags, UGC-style narration, 5 / 10 / 30 / 60 / 120 seconds, 3 credits.

    A substitute for lipsync when we see the mouth. Cheap VO makes the trap easier, not harder. The cheap version of a talking head is still two rows: this file, then a mouth pass.

    What is the VO sitting on?

    Look at the frame. Is there a mouth that should form these words?

    If no — pack shot, screen, B-roll, faceless UGC — Chatterbox TTS at 5s or 10s is the right 3-credit file. Tag the laugh. Stay inside 120s.

    If yes — creator talking to camera — generate the file if you want this voice, then lipsync. Do not lay 30s of tagged speech on a closed mouth and call it a talking ad.

    FAQ

    Is Chatterbox TTS a talking-head model?

    No. Text-to-audio, 3 credits, durations 5s / 10s / 30s / 60s / 120s, inline tags such as <laugh> and <sigh>. You get a voice file. A mouth on screen is a later row.

    How is this different from a TTS row with no duration list?

    This slug publishes a duration ladder. Length is a stop you pick, not only a script you hope is short. 120s is the ceiling. Other TTS rows govern by character cap instead. Do not paste a two-minute script into a 5s stop and blame the tags.

    Can I put <laugh> in the prompt of a video model instead?

    Not as a substitute for this row. The tags are documented on Chatterbox TTS. A video model may ignore them or invent a visual laugh you cannot use. Write the performance here; move the mouth elsewhere if we see it.

    Does it need a face upload?

    No. Nothing visual is in scope. If you have a face that must speak, that upload belongs on image-to-lipsync, with this 3-credit file as the track — or on a talking generate that is not Chatterbox.