Guides

    Seed Audio 1.0: a voice file, not a talking generate (4cr)

    Seed Audio 1.0 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Versely Team4 min read

    Seed Audio 1.0 writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.

    Seed Audio 1.0 is ByteDance text-to-speech: high-quality, natural-sounding TTS with preset voices, optional reference audio (@Audio1@Audio3) or a reference image, and speed, pitch, and volume controls. Category is text-to-audio. Content type is audio. Four credits is the catalog call. There is no resolution. There is no clip length enum. You get a voice file.

    Four credits for a wav, not a face

    This row does not move lips. It does not invent a presenter. It does not open a mouth that was generated closed. If the shot is a person talking on camera, Seed Audio 1.0 is the wrong first generate. A talking or lipsync model owns the mouth and the voice together. Laying this file on a silent face is how you get chewing on the wrong vowel.

    Off-screen narration, a bed under b-roll, a founder line over hands — that is this row. The AI text-to-speech tool is the door. Native audio versus TTS is the routing rule when a video model could have spoken in-shot instead.

    audio is true because the output is audio. That is not native-in-video speech. Do not confuse a TTS flag with a talking-head row.

    Presets, three audio refs, or a picture of the speaker

    The catalog description names three input shapes:

    • Preset voices. Pick one. Stop shopping mid-batch.
    • Reference audio @Audio1@Audio3. Up to three clips that steer timbre. Clean speech, not a song with a vocal buried in it.
    • A reference image. Optional. A face still is not a clone licence and it is not lipsync. It is a steer on the voice, not a presenter file.

    Speed, pitch, and volume are controls on the file you already wrote. They are not a reason to regenerate the words. Get the script right. Then ride the controls. Regenerating because the read was 5% fast is how four credits become forty.

    There is no duration slider because this is not video. Length follows the script. Write the line you will actually use. A paragraph you will cut in the timeline is how TTS starts to flatten; narration that goes long is a known failure, not a Seed Audio bug.

    Closed mouths stay closed

    If we can see the mouth, do not lay Seed Audio 1.0 on the take and call it done. Options that actually own the face:

    The failure that looks fine in headphones: you generate a cinematic clip, the mouth is idle or chewing, you drop Seed Audio on top, you ship. Viewers will not name the artefact. They will feel it.

    ByteDance's video family is on the ByteDance provider roster. Seed Audio 1.0 is the voice row, 4 credits, text-to-audio, presets plus @Audio1@Audio3 plus an optional image. The ranked TTS list is best text-to-speech model.

    FAQ

    Can I use Seed Audio 1.0 as the voice inside a talking-head generate?

    Not as a substitute for a talking or lipsync row. This model writes a file. If the mouth is in frame, drive the mouth from the file (lipsync) or pick a native-audio video model and write the line into that prompt. Do not stack.

    What are @Audio1@Audio3?

    Optional reference audio inputs named in the catalog description, up to three. They steer the voice. They are not a video in. They are not a lipsync driver unless you take the resulting TTS file to a lipsync row.

    Does a reference image make this an avatar?

    No. The image is an optional steer on the TTS. Content type is still audio. A talking presenter is a different category. Avatar X is one such row if the deliverable is a stock presenter.

    What do 4 credits buy?

    One Seed Audio 1.0 generate as listed: a voice file, text-to-audio, four credits. Not a clip, not captions, not a mouth. Script the line, pick the voice path (preset, audio refs, or image), set speed/pitch/volume after the words are right.