Seed Audio 1.0: One Model for Voice, Music, and Foley
ByteDance's Seed Audio 1.0 collapses TTS, music, and foley into one model with zero-shot cloning. How it works, and practical patterns for using it on Versely.
For most of the last two years, a finished piece of audio meant three separate tools: a TTS model for the voice, a music model for the score, and something else entirely — usually a sample library — for foley and ambience. Getting them to sit together meant a manual mix pass: level the dialogue, duck the music under it, loop the ambience so it doesn't repeat obviously, and hope the voice doesn't drift in tone across a longer script stitched from multiple generations. ByteDance's Seed Audio 1.0 is the first release that treats all three as one job instead of three handoffs, and it's worth understanding exactly where it earns that claim and where a dedicated tool still wins.
What Seed Audio 1.0 actually is
ByteDance launched Seed Audio 1.0 on June 23, 2026 at Volcano Engine's FORCE 2026 conference, positioning it as a unified audio generation model rather than another TTS release. Per that launch coverage, it generates natural speech, original music, foley effects, and environmental soundscapes from a single text description, with zero-shot voice cloning from a short reference clip.
The "unified" framing isn't just marketing. Most audio pipelines chain models together — script goes to a TTS engine, the output gets dropped into a DAW, ambience gets layered from a separate library. Seed Audio 1.0 generates complete audio scenes in a single pass: multi-character dialogue, sound effects, ambience, and music together, rather than one stem at a time.
What it actually does in one pass
Three capabilities matter for practical production work:
Multi-character dialogue with distinct voices. Seed Audio 1.0 supports multi-character dialogue in a single pass, assigning a different voice to each speaker rather than requiring separate generations stitched together afterward. Voice reporting on the model describes a structured dialogue timeline with 100ms timing precision underneath this — which is why lines land where you place them instead of drifting relative to each other across a longer script.
Zero-shot voice cloning. Feed it a short reference clip, and it carries that voice into new lines without a training step — the same zero-shot pattern that's become standard for cloning, applied here across dialogue, not just single-speaker narration.
Up to two minutes per request, with continuation. Generation runs up to two minutes of audio per request, with continuation support that keeps voice and ambience consistent when a piece needs to run longer than a single pass allows — so a three-minute scene doesn't mean three separately-cloned voices drifting apart.
Pattern 1: podcast-style two-voice dialogue
The clearest use case is a scripted two-person conversation — an intro-style cold open, a fake interview clip, a back-and-forth product explainer. The old way was two separate TTS generations, manually interleaved and timed. Seed Audio 1.0 takes the whole exchange as one script:
Generate a 45-second two-speaker dialogue.
Speaker A (warm, mid-30s female voice): "Okay, so everyone's
asking about the new pricing — walk me through it."
Speaker B (measured, older male voice): "Sure. The short version
is we simplified three tiers into one."
[continue exchange...]
Output: natural pacing, slight overlap avoided, studio-clean.
Model: Seed Audio 1.0.
That single request replaces what used to be two generations plus a manual edit pass to interleave them. On Versely, this pairs directly with the multi-speaker dialogue workflow for laying the result onto video with per-speaker timing already intact — useful for anything built on the voice-over side of a production rather than a single narrator track.
Pattern 2: ambience beds under b-roll
The second pattern is quieter but shows up constantly in faceless and UGC-adjacent production: a bed of environmental sound under silent or lightly-scored b-roll — café chatter under a coffee shot, street noise under a walking sequence, rain under a moody product spot. Rather than hunting a sample library for a clip that's close enough and looping it awkwardly, describe the environment directly:
Generate 30 seconds of ambient soundscape: light rain on a
window, distant traffic, occasional far-off thunder. No music,
no dialogue. Loop-friendly, consistent volume throughout.
Model: Seed Audio 1.0.
Because the model handles foley and environmental sound as a first-class output rather than a TTS side effect, this doesn't require pretending a speech model can also do ambience — it's the same generation surface, just a different kind of request.
Why single-pass generation matters beyond convenience
The bigger win isn't the time saved chaining fewer tools — it's consistency. When dialogue, ambience, and score come from three separate generations, each one has its own implicit "room tone," and a listener registers the seams even if they can't name what's wrong. A voice recorded against silence and a voice recorded against a synthetic café bed don't sit together the same way a voice generated with that café bed in the same pass does. Single-pass generation doesn't guarantee broadcast-quality mixing, but it removes the specific failure mode where each layer sounds like it was captured somewhere slightly different — the audio equivalent of a video shoot where every clip was lit and mic'd on a different day.
When a dedicated music model still wins
Seed Audio 1.0's music generation is part of the same unified pass, which makes it the right call for a short score or an ambient bed under a scene. It's not the right call when the job is a structured song — verse, chorus, a hook that needs to land at a specific bar. For that, a purpose-built music model with more control over song structure, like Versely's Suno-based options, is still the better tool. The practical split: Seed Audio 1.0 for anything that's dialogue-plus-environment or a simple mood bed, a dedicated music model for anything that needs to function as an actual song.
Pricing on Versely
Seed Audio 1.0 is priced at 4 credits per 1,000 characters of script, billed on the text you send in regardless of how many speakers or how much ambience the request includes — a two-voice dialogue and a single-narrator script of the same length cost the same.
FAQ
Does Seed Audio 1.0 replace a dedicated text-to-speech model?
For most narration and dialogue work, yes — it covers standard TTS plus voice cloning in the same request. Versely still carries several dedicated TTS engines for cases where a specific voice character or language coverage matters more than the unified-audio convenience.
Can it generate a full song with lyrics?
It generates music as part of its unified output, which is well suited to scores and ambient beds, but a dedicated text-to-music model is the better choice when the job is a structured song with verses and a chorus rather than a mood track.
How long can one generation run?
Up to roughly two minutes per request, with continuation support for pieces that need to run longer while keeping the same voices and ambience consistent across the extension.
What do I need to clone a voice?
A short reference clip is enough — the cloning is zero-shot, meaning no separate training step before you can generate new lines in that voice.
Does it handle sound effects, or only ambience beds?
Both. Foley — discrete effects tied to an action, like a door closing or a glass set down — and continuous environmental soundscapes are the same category of output for the model; the difference is whether your prompt describes a single event or a sustained scene.
Try the two-voice dialogue pattern first — it's the fastest way to feel the difference between "one model generating a scene" and "two TTS calls stitched together." Versely's AI text-to-speech tool has Seed Audio 1.0 available alongside the rest of the catalog if a specific job calls for a different engine instead.