The format has hard conventions and they are all borrowed from the feed it hides in: a face talking directly to the lens, the product physically in hand, cut-ins to the thing being described, burned-in captions, no title sequence and no music bed doing emotional work. Break enough of them and the piece stops passing as a post, which is the only advantage the format has.
That makes it the one place in generated video where the failure mode is looking too good. A cinematic grade, an impossible camera move or a studio-lit product shot all read as advertising instantly. The brief is closer to "believable phone footage" than to anything the model was rewarded for during training, so it usually has to be asked for explicitly.
In Versely a UGC piece is assembled rather than generated in one shot: a talking-head clip laid over the product footage, reaction and b-roll clips from the hooks library, an avatar and voice when nobody is going on camera, and captions burned in at the end. Several of the public workflow templates — the influencer review, the panel reaction reel — are this shape already, which is usually a faster start than a blank prompt.
In practice
- Show the product being used, not presented — hands in frame beat a clean pack shot.
- Ask for phone-camera imperfection explicitly: slight handheld motion, ordinary indoor light, no colour grade.
- Captions are part of the format, not a finishing touch — the format is watched with the sound off.
Avatar models
Catalog entries that ship with ready-made presenters. 4 of the 296 models in the Versely catalog qualify.
| Model | Provider | Type |
|---|---|---|
| HeyGen Avatar V5 | HeyGen | Lipsync |
| VEED Avatars | VEED | Lipsync |
| HeyGen Avatar V3 | HeyGen | Lipsync |
The mistake to avoid
Producing one perfect UGC piece. The format's economics come from volume and variation — several creators, several angles, several openings — and a single polished exemplar is the version least likely to look like a post.
Where you will run into it
- Add Picture-in-Picture to a Video — Reaction footage, corner-mounted on your main clip.
- Add Reaction Clips to a Video — Genuine reaction footage, dropped into your cut.
- AI UGC Video Generator — UGC ads at the speed and price of a prompt.
- AI Testimonial Video Generator — The words are already yours. This gives them a face.
- AI Avatar Generator — Pick a face, hand it a script, get a presenter.
- Plush Lipstick Influencer Review — Sara the influencer reviews the Plush Baby Pink Lipstick in 3 scenes: greeting, product review with features, and promo code CTA. Consistent voiceover narration throughout.
- Viral Panel & Reaction Reel — A 4-scene viral hook video. Scene 1-2: A professional panel expert discusses her before/after results with a product, denying she's gatekeeping. Scene 3-4: An influencer films a selfie-POV reaction, holding the product and confirming the expert's claims.
Related terms
Hook rate
Hook rate is the share of people shown a video who are still watching a few seconds in — the number that grades the opening, not the edit behind it.
Creative fatigue
Creative fatigue is performance decaying because the audience has seen the same piece too many times, rather than because anything about the piece changed.
B-roll
B-roll is the supporting footage cut over narration or an interview — everything on screen that is not the person doing the talking.
Brand kit
A brand kit is the saved set of colours, fonts, logo, product images, tone of voice and default frame shape that generations are expected to obey without being reminded.
Burned-in captions
Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.