Guides

    Grok TTS: a voice file, not a talking generate (4cr)

    Grok TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team5 min read

    Grok TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Grok TTS is xAI's text-to-audio row: high-quality speech with 5 expressive voices, 21 languages, and speech-tag support. Up to 15,000 characters per request. The catalog lists 4 credits. Audio is on because audio is the product. No image is required. There is no duration ladder and no resolution — this file has no frame. Category is text-to-audio only.

    That is a voice file. It is not a presenter. Laying it under a still whose mouth never moves is fine. Laying it under a face that should be saying the line is how you ship a puppet.

    Five voices is a cast, not a face

    Five expressive voices and twenty-one languages is a lot of range for 4 credits. It is still a cast list for sound. Speech tags belong in the script — the breaths, the emphasis, the asides the voice should perform. They do not open a jaw.

    If the video is faceless — screen recording, pack orbit, B-roll — Grok TTS is the pass. Generate the VO, drop it on the timeline, caption it. The text-to-speech tool is the launcher. Add voiceover is the assembly step.

    If we see the mouth, stop. A talking or lipsync row takes a plate plus this kind of track (or a script of its own). Grok TTS can still write the track. It cannot be the only model in that job. The ranked speech map is best text-to-speech. The talking-head map is best lipsync.

    Fifteen thousand characters is a script

    The cap is 15,000 characters per request. That is an article, a long VO, a localisation pass — not a five-second hook you should have typed in one sitting. Use the cap when the job is long. Do not paste 15,000 characters to "see what happens" on a bumper that needed two sentences.

    There are no 5s / 10s / 30s buttons on this row. Length follows the script. If you need a five-second read, write a five-second script. If you need a two-minute read, write it and stay under 15,000 characters. Other TTS rows in the catalog expose duration steps; Grok TTS does not. The character cap is the governor.

    Speech tags are supported. Use them as performance notes in the text, not as a way to fake a visual laugh. A tagged laugh on a closed mouth is still a closed mouth with a laugh in the speakers.

    What this row will not do

    It will not generate a clip. Content type is audio. If you only run Grok TTS, you have a file you can play in the dark.

    It will not lipsync. There is no face input. requires_image is false because there is nothing to show.

    It will not replace Grok's video rows. Imagine is a different job on the same roster. Mixing those labels is how a 4-credit voice file gets asked to "make the person talk."

    Credits: 4 on the listed row. That is cheap next to a talking generate, which is the point — and also the trap. Cheap VO on a talking picture that never moved is still a failed talking picture. Spend the 4 credits on the file. Spend a lipsync row on the mouth. Do not split the difference.

    Is there a mouth in the cut?

    Look at the frame. Should it be forming these words?

    If no, Grok TTS is the right 4-credit file. Pick a voice, pick a language, tag the performance, stay under 15,000 characters.

    If yes, Grok TTS may still write the track, but the generate you need is lipsync or a talking model. A closed mouth with perfect xAI speech is the uncanny version of "we forgot the picture pass."

    FAQ

    Is Grok TTS a talking-head model?

    No. It is text-to-audio. Five voices, twenty-one languages, speech tags, 15,000-character requests. You get a voice file listed at 4 credits. A mouth on screen is a different row.

    Can I use it as the audio input for lipsync?

    Yes — that is a two-row job. Write the track here, then take a plate plus the file to a lipsync model. Do not skip the second row and hope the editor will "feel" like talking.

    Why is there no duration picker?

    The catalog does not list durations on Grok TTS. Length is the script, capped at 15,000 characters. If you need a timed bumper, write to time. If you need a duration-step TTS row, that is a different catalog line.

    Does it need a reference image?

    No. Nothing in the frame is this model's problem. Upload a face only when you have moved to the lipsync pass. On this row an image is unused budget in the wrong slot.