The roster is the product. You choose a presenter who already exists in the catalog — a stock face, often with a stock voice — and you never photographed them. That is a different door from uploading a portrait, and a different door from driving a mouth on a clip you shot.
Lipsync is the mouth-matching job and will run on a photo you bring; it does not invent the identity. Character-consistency asks whether that identity holds across later jobs. UGC is the ad format that often wants a talking face. A premade avatar answers a prior question: whose face is this, and did you supply it.
Versely's avatar generator splits the two routes on purpose. The agent can list the roster first; a photo-driven talking-head is a different request even when both end up as a presenter clip.
In practice
- Pick from the roster when you need a presenter and have no consented face to upload.
- Treat the chosen avatar as the series identity — swapping stock presenters mid-run reads as a recast.
- A photo-in talking head is a different job; do not send a portrait into a roster picker and expect it to be used.
Avatar models
Catalog entries that ship with ready-made presenters. 4 of the 305 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| HeyGen Avatar V5 | HeyGen | Lipsync |
| VEED Avatars | VEED | Lipsync |
| HeyGen Avatar V3 | HeyGen | Lipsync |
The mistake to avoid
Using a premade avatar when the brief is that a specific real person is speaking. The roster is identity you did not supply, and the audience can tell.
Where you will run into it
- AI Avatar Generator — Pick a face, hand it a script, get a presenter.
Related terms
Lipsync
Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.
Character consistency
Character consistency is whether the same person, mascot or product still looks like itself across separate generations.
UGC
UGC is content made by an ordinary person — a customer, fan or creator — rather than by the brand itself, and it now covers everything from an unprompted review to paid creator work built to look exactly like one.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.