Guides

    Gemini 3.1 Flash TTS: a voice file, not a talking generate (4cr)

    Gemini 3.1 Flash TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team4 min read

    Gemini 3.1 Flash TTS writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Gemini 3.1 Flash TTS is Google's text-to-audio row: expressive speech with 30 voices, natural-language style control, inline audio tags ([sigh], [laughing], [whispering]), multilingual synthesis, and multi-speaker dialogue. The catalog type is audio. It does not take an image. It does not list a duration chip or a resolution. The display price is 4 credits. The audio flag on the row is false because this is not a talking clip — it is a voice file.

    Flash TTS is a wav, not a face

    You type. You get speech. That is the whole job. Style control and the 30-voice roster live in the script: which speaker, how they say it, which tag punches the breath. Multi-speaker dialogue is still a soundtrack. Two names in the prompt do not become two faces. They become two voices on a file you can drop under a cut that already exists.

    The AI text-to-speech tool is the door for this row. The Google roster also has Veo and Gemini video. Those are different content types. Do not infer a presenter because the same vendor ships one.

    If the brief is "we need the line," this is the row. If the brief is "we need the mouth," you left this row.

    Tags steer delivery, not the plate

    Inline tags are the feature people quote and then misuse. [whispering] changes how the take is performed. It does not open a jaw. Natural-language style control is the same class of lever: warmer, slower, a second speaker entering. Useful. Still audio.

    Long blocks will flatten even with a tag at the top. Chunk at the sentence. That is why narration goes flat, and it is true on Flash TTS the same as on any other speech engine. Audio tags punch a moment. They do not typeset a face.

    Multi-speaker is the other trap. A two-hander in the script is two performances on one file. If each speaker needs a picture, you still need plates and a lipsync pass after this file exists.

    When we see the mouth, change the row

    A closed mouth with this wav laid under it is the tell. You paid 4 credits for speech and then asked the picture to pretend. That is not a Flash TTS miss. That is a routing miss.

    The talking path is a lipsync or talking-head row that requires a plate and a track. Kling Avatar Pro is one: image in, audio in, 4K talking avatar out, 58 credits. Run it from AI lipsync. Generate the voice here (or on another speech row), lock the still on an image row, then do the mouth pass. Do not ask Flash TTS to skip those steps.

    The ranked speech list is best text-to-speech model. Flash TTS sits on it as a 4-credit Google file. It does not sit on the talking-head list.

    Four credits, then stop

    There is no 5s/10s slider because there is no clip. There is no 720p or 4K because there is no frame. You are billed for the speech job at 4 credits. That is cheap enough that teams burn it on "maybe we'll put a face on it later." Later is a different model. Budget the 4 credits as the track. Budget the picture and the mouth separately.

    If you only needed the line — VO under b-roll, a phone-hold, a cutaway — you are done when the wav lands. Do not "upgrade" this generate into a talking clip by hoping. Switch the row.

    FAQ

    Does Gemini 3.1 Flash TTS make a talking video?

    No. It is text-to-audio. You get a voice file: 30 voices, style control, tags, multilingual, multi-speaker. The mouth is a lipsync or talking-head row that wants a plate.

    Why is the audio flag false on a TTS model?

    On this catalog, audio: false means the row is not a video generate with a soundtrack. Flash TTS is the soundtrack. Treat it as a file you attach, not a clip that already has a face.

    Can I lay this wav on a still and call it a presenter?

    You can attach it. The mouth will not track unless you run a lipsync model after. If we see the mouth, pick that row instead of arguing with a closed plate.

    Is 4 credits the whole cost of a talking ad?

    It is the cost of this speech pass. The still and the mouth pass are other rows with other credit numbers. Plan those. Do not hide them inside Flash TTS.