The roster is the product. You choose a presenter who already exists in the catalog — a stock face, often with a stock voice — and you never photographed them. That is a different door from uploading a portrait, and a different door from driving a mouth on a clip you shot.
Lipsync is the mouth-matching job and will run on a photo you bring; it does not invent the identity. Character-consistency asks whether that identity holds across later jobs. UGC is the ad format that often wants a talking face. A premade avatar answers a prior question: whose face is this, and did you supply it.
Versely's avatar generator splits the two routes on purpose. The agent can list the roster first; a photo-driven talking-head is a different request even when both end up as a presenter clip.
In practice
- Pick from the roster when you need a presenter and have no consented face to upload.
- Treat the chosen avatar as the series identity — swapping stock presenters mid-run reads as a recast.
- A photo-in talking head is a different job; do not send a portrait into a roster picker and expect it to be used.
Avatar models
Catalog entries that ship with ready-made presenters. 4 of the 331 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| HeyGen Avatar V5 | HeyGen | Lipsync |
| VEED Avatars | VEED | Lipsync |
| HeyGen Avatar V3 | HeyGen | Lipsync |
The mistake to avoid
Using a premade avatar when the brief is that a specific real person is speaking. The roster is identity you did not supply, and the audience can tell.
Where you will run into it
- AI Avatar Generator — Pick a face, hand it a script, get a presenter.
Related terms
Lipsync
Lipsync meaning: an AI model moves a face's mouth to match an audio track, so the words look spoken by that face instead of dubbed over the top.
Character consistency
Character consistency meaning: whether the same person, mascot, or product still looks like itself across separate generations.
UGC
UGC meaning: user-generated content from customers or creators, not the brand. In 2026 it also covers paid UGC-style ads that look native.
Voice cloning
Voice cloning meaning: building a reusable synthetic voice from a real sample so new scripts can be spoken in that voice later.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.
Free on iPhone. On a computer? The same account works on Versely Web, no install needed.