Qwen 3 TTS Voice Design: a voice file, not a talking generate (3cr)
Qwen 3 TTS Voice Design writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Qwen 3 TTS Voice Design writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Qwen 3 TTS Voice Design is Qwen's audio row: text-to-audio and voice-clone, 3 credits, no picture. The catalog line is "Qwen 3 text-to-speech with custom voice design." You describe the voice, then you give it words. You do not get a jaw. The text-to-speech tool is the door.
Design the voice, then speak the line
Voice design is not a talking generate. It is a brief for how the file should sound: age, grit, pace, room. Keep that brief out of the script field. The script field is the words. Mixing a character bible into the spoken line is how the model says the bible. Text-to-speech is the mode. Voice cloning is listed as a category on this record — still audio, still not a mouth.
Three credits is the catalog price. There is no resolution and no duration ladder. This is not a 4K avatar. The Qwen roster is the rest of the brand. This page is only the TTS voice-design row.
If we see the mouth, this file is incomplete
A designed voice under b-roll is a finished VO. A designed voice under a silent face is a dub you have not admitted yet. Run AI lipsync or a talking row when the jaw is in frame. Add a voiceover when it is not. Do not hope the waveform will invent lips.
If you need a specific real person, an upload-a-sample clone may be the better category cousin. If you need a voice that does not exist yet, this is the row: describe it. Either way, the output is a file you can lay or a file you can lipsync. It is never the scene.
Two fields, one file
A useful Qwen 3 pass is a voice description plus a clean script. Custom voice design is the whole point of the name. Do not skip the design and then blame the model for a generic read. Do not skip the mouth decision and then blame the model for a closed jaw.
The best text-to-speech list is the wider audio map. Stay on this row when the job is "invent the voice, speak the line, stop." Three credits. Audio out. Picture elsewhere.
FAQ
Does Qwen 3 TTS Voice Design make video?
No. Content type is audio. Categories are text-to-audio and voice-clone. You get a designed voice file, not a talking clip.
Is "voice-clone" the same as uploading a sample?
This record is custom voice design — you describe the voice. Treat it as a design brief, not as a guarantee that a real person's sample is the input. The output is still audio only.
Do I need an image?
No. This row does not require an image. An image will not open a mouth. Use lipsync when we see the jaw.
What are the 3 credits for?
The catalog price for this text-to-speech generate. Not a scene, not 4K, not a duration.