Qwen 3 TTS 0.6B: a voice file, not a talking generate (2cr)
Qwen 3 TTS 0.6B writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Qwen 3 TTS 0.6B writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
Qwen 3 TTS 0.6B is a compact text-to-speech model with natural voice synthesis and efficient processing. Category is text-to-audio. Two credits. Content type is audio. audio is true because the output is speech, not because a video was generated. Features: text-to-speech, multilingual, natural voice. No resolution. No duration enum. You get a small, cheap voice file.
Two credits is the compact row
Inworld 1.5 Max is 4 credits with a 15-language library. Seed Audio 1.0 is 4 credits with @Audio1–@Audio3 and pitch controls. Qwen 3 TTS 0.6B is 2 credits, compact, efficient, natural. That is the batch VO, the scratch read, the localisation draft you can afford to throw away. It is still not a face.
Efficient processing is a listed feature, not a licence to skip review. Compact models flatten on long paragraphs. Write short lines. Narration that goes long is the failure mode. Two credits per call is how you split a script instead of feeding it a monologue.
The AI text-to-speech tool is the door. Qwen's provider roster is the family. Best text-to-speech is the ranked list. This page is only 0.6B: 2 credits, compact, multilingual, audio file out.
Mood is a parameter. Do not tag the script.
Qwen sits with the providers that take emotion as a separate field, not as bracketed markup in the words. Put tags in this script and they get read aloud. Clean text in the script field. Mood in the style field. That split is the whole prompting rule for this slug.
Multilingual is listed. It is not "15 languages" and it is not Inworld's voice library. Do not copy another TTS page's language count onto 0.6B. Use it for natural reads across languages the row actually serves, then listen. Compact is not Max.
There is no clip length slider. Length follows the script. There is no 720p because there is no picture.
Compact is not a presenter
Two credits is how people justify laying Qwen on everything, including a silent cinematic plate with a face in close-up. Cheap does not fix desync. If we see the mouth:
- Native-audio video model, line in the prompt, or
- Lipsync driven from this 2-credit file, on purpose
If we do not see the mouth: Qwen, then add voiceover. Captions after the read is locked. Native audio versus TTS is the brief-level rule.
Do not pick 0.6B as a talking-head because the catalog says audio: true. That flag means the TTS emits sound. It does not mean a mouth was generated.
FAQ
Is Qwen 3 TTS 0.6B a video model?
No. Content type is audio. Two credits buys a compact voice file. A talking generate is a different category.
Why 2 credits when other TTS rows are 4?
That is the catalog figure for this compact 0.6B row. Efficient processing is listed. The job is the same: a file, not a face. Do not treat cheap as "close enough for a close-up."
Can I put delivery tags like [whisper] in the script?
Not on Qwen. Mood belongs in the separate style/emotion field. Markup in the script is likely to be spoken as text. Clean words, then a style instruction.
What do 2 credits buy?
One Qwen 3 TTS 0.6B generate: compact multilingual TTS, natural voice, 2 credits. Not a clip, not captions, not lipsync. If the mouth is in frame, take this file to a lipsync row or pick a talking model instead.