Guides

    Avatar X Text to Video: the mouth pass after the picture exists (720p, 150cr)

    Avatar X Text to Video wants a plate and a track. It is not how you invent the scene.

    Versely Team4 min read

    Avatar X Text to Video wants a plate and a track. It is not how you invent the scene.

    Avatar X Text to Video is the script-to-presenter row: type a script and a stock avatar performs it — voice included. Mirage's most advanced avatar model, with strong identity preservation and expressive delivery. Scripts run 50–1,500 characters. Category is text-to-lipsync. Content type is lipsync. audio is true. 720p. 150 credits is the catalog figure, billed per second, with a 150-credit minimum. requires_image is false because the plate is a stock avatar, not a world you described.

    150 credits, 720p, a person talking

    The file is a talking presenter. The face, the voice, and the read are the file. A product turntable sent here comes back as a person holding something that is not your product. A landscape brief sent here comes back as a host in front of a backdrop the model invented. Both are the same routing error: you asked a lipsync row to invent a scene.

    Stock avatars are the plate. The script is the track. Voice is included. You do not write a cinematic paragraph. You write 50 to 1,500 characters of speech — one thought per take. Under 50 is not a script. Over 1,500 is a monologue this row will not take.

    The AI lipsync tool is the assignment door. The AI avatar generator is the presenter surface. Best talking-head model ranks the category. A longer version of the presenter test is Avatar X when the presenter is the file. This page is only the text-to-video slug: 720p, 150 credits, 50–1,500 characters, voice included.

    The scene is not the job

    There is no duration enum because length follows the script. There is no 21:9 cinematic ratio on this row. Quality is 720p. That is enough for a medium close-up. It is not a reason to pick Seedance instead "for resolution." Upscale after picture lock if the placement needs it. Do not buy a world generator to get 1080p you will crop to a face.

    requires_image is false. You are not required to upload a custom face. Stock roster first. A custom likeness is a different input shape and a consent problem, not a prompt adjective.

    Do not ask Avatar X to typeset a lower third. Generate the presenter, lock the take, then burn captions. On-screen type inside a lipsync generate is how you get misspelt titles on a face you otherwise liked.

    Mouth pass, not world pass

    If the picture does not exist yet — no avatar picked, no idea who is speaking — this row will still run, because stock avatars exist. That is not "invent the scene." That is pick a roster face. The scene around them is a backdrop. If you needed the factory, the wet bar, or the weather, generate that on a scene model and cut. Do not spend 150 credits a second (matrix max 3600) hoping the avatar becomes a location scout.

    150 credits is the listed minimum. Talking-head work is expensive relative to a 12-credit silent Turbo plate. That is the point: you are paying for identity hold and a read, not for motion B-roll. Call this the hero only when the deliverable is a person saying a script.

    If the voice must be a cloned founder, this row's included voice is not that contract. TTS or clone first, then a lipsync path that can take the track. Avatar X Text to Video is script in, stock avatar + voice out.

    FAQ

    Can I use Avatar X Text to Video to generate a product video with no presenter?

    No. The catalog job is a stock avatar performing a script, voice included, 720p. A product orbit is a different category. You will get a person. You will not get your pack shot.

    Do I have to upload a photo of the speaker?

    No. requires_image is false. Text-to-video on this slug runs on premade avatars. A reference-face pipeline is a different input shape.

    Why 720p when other video rows list 1080p or 4K?

    Because the file is a talking head. max_output_resolution is 720p. Upscale after lock if the placement needs it. Do not switch to a scene model just to buy resolution you will crop away.

    What do 150 credits buy?

    The catalog figure is 150, billed per second, minimum 150. You are buying a lipsync generate: 50–1,500 characters of script, stock avatar, voice included, 720p. Not a world, not a turntable, not burned-in captions.