MiniMax Speech: a voice file, not a talking generate (5s, 10s, 30s, 2cr)
MiniMax Speech writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth.
MiniMax Speech writes audio. If we see the mouth, pick a talking or lipsync row instead of laying this on a closed mouth. The catalog description: “MiniMax's text-to-speech engine — clear multilingual delivery backed by one of the largest named-voice rosters in the catalog, so you can cast a specific voice rather than settle for a generic one.” Category is text-to-speech / text-to-audio. Audio is on. No image required. Durations listed: 5s, 10s, 30s, 60s. Catalog credits: 2, marked discounted in the snapshot.
MiniMax Speech returns a voice file. It does not open a jaw. It does not invent a presenter. If the shot is a face that must speak, generate or lock the picture first, then either pick a talking video row or run lipsync on the plate with this track. Dropping a MiniMax read under a still mouth is how you get a radio ad on a photograph.
Named voices are the product. Seconds are the file length.
The pitch is the roster: cast a specific voice, multilingual, clear. That is a different job from “a video model that happens to talk.” Audition the same line in two named voices. Keep the one that hits the consonant. Do not pick MiniMax Speech because a Hailuo video demo felt warm. Hailuo is MiniMax’s video line on a different roster. This row is speech.
Duration options are 5s, 10s, 30s, and 60s. That is how long the file can be, not a mouth-sync guarantee. A 60-second read is a voiceover. A 5-second read is a bumper. The price matrix bills per 1,000 characters, not per second of playback — cutting words is how you cut the bill. The 2-credit sticker (discounted in the catalog) is the row’s list figure, not a per-second video meter.
The text-to-speech tool is the door. The ranked list is best text-to-speech model. Use MiniMax Speech when you need a castable named voice as a file. Use a video model when you need a picture.
If we see the mouth, this row is the track, not the picture
Two honest pipelines:
- Voiceover on B-roll. Generate or shoot the picture with a closed mouth or no face. Lay the MiniMax file under it. Add voiceover to video is the edit task. This is the MiniMax-shaped job.
- Talking face. Lock the plate. Generate the read. Run a lipsync row so the jaw matches the track. Do not paste the wav on a still cheek and hope.
A third, worse pipeline is asking a native-audio video model to “also be MiniMax.” It will not be. It will cast its own bed. If you already paid for a named MiniMax voice, protect it: mute the video generate or replace the bed.
Script like speech, not like a caption. Named voices expose stiff copy. Read it out loud. If you would not say it, do not send it. Multilingual delivery is in the description — name the language, do not assume a later caption pass is a dub.
Two credits is not a video budget
Teams see 2 credits (discounted) and try to “make the ad” on this row. You will get a beautiful wav and no picture. That is success for TTS and failure for a talking generate. Buy the stills row. Buy the motion row. Buy this row for the voice. Three rows, three contracts.
There is no resolution because there is no picture. There is no 1080p, no 4K, no aspect ratio list. If your brief includes those words, you are not on MiniMax Speech yet. Come back when you need the file that comes out of a speaker.
FAQ
Will MiniMax Speech animate a face?
No. Content type is audio. Category is text-to-audio. For a moving mouth, use a talking generate or a lipsync row with this file as the track.
What lengths are listed?
5s, 10s, 30s, and 60s. A minute is on the menu. That is still a voice file.
Does it need an image?
No. requires_image is false. An image means you wandered into a video or lipsync row.
Why is it marked discounted?
The catalog snapshot flags this row discounted, with a 2-credit list figure. That does not make it a cheap video model. It makes a named-voice wav inexpensive relative to a 70-credit talking generate — which is the point, if you actually want a wav.