AI Models

    Avatar X when the presenter is the file

    Catalog slug avatar-x. Use it when the deliverable is a talking presenter, not a generated world.

    Versely Team6 min read

    A product turntable sent to Avatar X comes back as a talking person holding something that is not your product. A presenter brief sent to Seedance comes back as a world with someone standing in it who is not a presenter. Both are the same routing error: you asked for a file the model is not built to emit.

    Avatar X exists for one job. The deliverable is a talking presenter. The face, the voice and the read are the file. If the brief is a generated room, a product orbit, weather, or a three-beat gag, this is the wrong row in the catalog.

    What Avatar X actually emits

    Avatar X is Mirage's talking-presenter model. You type a script and a stock avatar performs it, voice included. Scripts run 50 to 1,500 characters. Output is 720p. Identity hold is the feature: the same face, the same delivery character, episode after episode.

    There is a second input shape if you already have the audio. Avatar X Reference to Video takes a reference photo or clip plus 3 to 180 seconds of audio and drives the mouth from that track. Text-to-video is "write the line, get a presenter." Reference-to-video is "here is the voice, put it on this face." They are the same job family. They are not a world generator with a person bolted on.

    The AI avatar generator is the surface for that job: pick a roster face or lock one from a still, then drive it from a script or a track. That is a different pipeline from text-to-video models that invent a scene and happen to put a mouth in it.

    Burn captions after picture lock. The presenter file is silent of burned-in type until you add captions. Do not ask Avatar X to typeset the lower third.

    The presenter test

    Ask one question before you spend a credit: is the file a person talking, or a world in which talking happens?

    If the file is a person talking — a course host, a weekly explainer, an internal standup, a product FAQ read to camera — Avatar X is the model. The camera is a medium close-up. The set is a backdrop. The mouth is the shot. Identity has to survive a series, which is why you lock a face first and reuse it.

    If the file is a world in which talking happens — a founder walking a factory, a bottle on a wet bar, a sky that has to look like weather — you are not in the avatar row. Scene models invent rooms. They will invent a face to match the room. That face will not be your presenter next Tuesday.

    The test is about the deliverable, not about whether anyone speaks. Plenty of scene clips have dialogue. Dialogue in a world is still a world job. A talking presenter with no world worth keeping is an avatar job. Write the script as a script — 50 to 1,500 characters, one thought per take — and lock the face before you scale the series. Identity is the only reason to be on this row; if you re-roll a new stock avatar every episode you have already left the job.

    Do not send a product turntable here

    A turntable is an object rotating from angle A to angle B. The contract is the object: label, cap, seam, colour. Avatar X's contract is a face. Give it a bottle and a script and you get a person describing a bottle the model invented. The pack shot you needed is gone.

    Product rotation is a first-last-frame job. World continuity is a reference-to-video job on a scene model. Neither of those is what Avatar X is for, and burning presenter credits on them is how teams decide "avatars don't work" after one afternoon.

    The other direction is equally expensive. Do not send a presenter brief to Seedance. Seedance will give you a cinematic person in a cinematic room. That is a scene. It will not give you a reusable host. The next episode will be a cousin of the last one, which is the failure mode a stock avatar exists to prevent.

    When the presenter is not enough

    Avatar X does not own every talking shot.

    • The mouth is small in frame. If you cannot see the lips, you paid for lipsync you cannot use. A wide or a landscape with a voiceover is a scene-plus-audio job.
    • The product has to be the real product. A presenter holding a pack is not a pack shot. Shoot or generate the object separately and cut.
    • The voice has to be a cloned founder across thirty episodes. Avatar X generates a read with the avatar. A locked brand voice is a TTS or clone pass, then the reference-to-video route if the mouth must match that track.
    • Health, legal, finance, politics, presented as a human expert. A synthetic presenter in those categories is a YouTube inauthentic-content problem, not a model-quality problem. The file can be technically fine and still be a Partner Program kill.

    If the brief clears those traps and the deliverable is still "this person says this script," stay on Avatar X. Do not upgrade to a scene model because the stills look prettier. Prettier rooms are not the job.

    FAQ

    Is Avatar X a text-to-video model I should use for ads?

    Only if the ad is a talking presenter. UGC-style talking heads, course hosts, and FAQ reads are in scope. Product hero shots, lifestyle worlds, and turntables are not. Route those elsewhere and cut the presenter in if you need both.

    Do I need my own face to use it?

    No. Text-to-video runs on stock avatars. Reference-to-video can take a photo or a clip of a real face if you have consent. The avatar generator is built around that split: roster first, custom face when you have one.

    Why is output 720p if other models go to 1080p or 4K?

    Because the file is a talking head, not a landscape plate. 720p on a medium close-up is a different delivery problem from 4K on a wide. Upscale after picture lock if the placement needs it. Do not pick a scene model just to buy resolution you will crop away.

    Can I caption inside the generation?

    No. Generate the presenter, lock the take, then burn captions. Asking the avatar model to render on-screen type is how you get misspelt lower thirds on a face you otherwise liked.