Character Description Prompts for Consistent People
Character description prompts for consistent AI people: writing a character block, which traits stick between generations, and when to move to references.
Generate "a friendly barista" five times and you'll meet five strangers. That's the character consistency problem in one sentence, and it's the wall every serialized AI project hits — the brand mascot who shape-shifts between posts, the tutorial host who ages ten years between scenes, the story protagonist nobody can follow because she's a different person in every shot. Prompt text alone can't fully solve it, but it can get you 70% of the way, and knowing how to write a character description block — and when to stop relying on words and switch to reference images — is the actual skill. Here's both halves.
Write a character block, not a description
The core practice: write your character once, as a fixed block of text, and paste it verbatim into every prompt. Not similar wording — identical wording. Models are sensitive to phrasing, and "a woman with red curly hair" versus "a curly-haired redhead" can pull noticeably different faces from the distribution. Paraphrase is how characters die.
A working block is 25–45 words, structured like a casting sheet:
MAYA: a woman in her early 30s, warm brown skin, black box braids past her shoulders, round tortoiseshell glasses, small gold hoop earrings, wearing a rust-orange corduroy jacket over a cream tee.
Then every scene prompt becomes: block + action + setting. "MAYA: [block]. She laughs while pouring coffee in a sunlit kitchen, medium close-up." The name prefix does nothing magical for the model — but it does something for you: it keeps the block atomic in your prompt library, so nobody on the team "improves" three words of it in scene twelve.
Which traits actually stick
Not all description survives regeneration equally. Traits rank roughly like this:
| Stickiness | Traits | Notes |
|---|---|---|
| High | Wardrobe, glasses, distinctive hair (color + style), accessories, facial hair | Discrete, nameable, high-training-signal — the backbone of recognition |
| Medium | Age range, build, skin tone, hair length | Hold broadly, drift in degree |
| Low | Exact facial structure, "a kind face," subtle features | The actual face lottery — text cannot pin a specific stranger's bone structure |
The design consequence: build recognition out of high-stickiness traits. Audiences identify recurring characters by silhouette, hair, and wardrobe far more than by precise facial geometry — animation studios have exploited this forever. A character defined by box braids, tortoiseshell glasses, and a rust corduroy jacket survives face-lottery drift, because viewers track the braids and the jacket. A character defined as "a beautiful woman with an approachable smile" is unrecognizable by design.
Corollary: resist the urge to over-specify low-stickiness traits. Ten clauses about cheekbones and eye spacing don't pin the face; they just crowd out the wardrobe details that would have carried recognition.
The escalation ladder: text → image anchor → reference models
Text blocks are stage one. When the project demands tighter consistency, escalate deliberately:
Stage 1 — text block only. Good enough for: characters seen in different contexts days apart, stylized/illustrated characters, background figures. Cheapest, fastest, no asset management.
Stage 2 — image anchor + image-to-video. Generate (or photograph) one definitive portrait of your character, then drive video from that still. The image pins face, wardrobe, and palette absolutely for that clip; your prompt only carries motion. The catch: each clip starts from the same frame unless you build variations — the practical patterns for chaining this across a multi-scene sequence are covered in character consistency across scenes with an I2V fallback chain.
Stage 3 — reference-to-video models. The purpose-built answer: models that accept character reference images as a first-class input and then render that person in new scenes, angles, and actions described by your prompt. VEO 3.1's reference-to-video mode and Wan 2.7 reference-to-video both work this way — you supply the face once, and prompt text goes back to doing what it's good at: action, setting, camera. This is the stage for serialized content where the same host appears in every episode.
Stage 4 — avatar systems. When the character is a presenter who must speak to camera repeatedly — same person, new script, every week — the economics flip toward a dedicated AI avatar built once and driven by scripts thereafter.
The mistake is starting at stage 1 and white-knuckling it through a 20-scene project with re-rolls. Price the escalation early: if the character appears more than a handful of times, stages 2–3 cost less than the retake budget you'll otherwise burn.
Keeping the block alive across a real project
Consistency decays through operational sloppiness more than model limits. Field rules:
- Version the block like code. If Maya's jacket changes for a winter campaign, that's MAYA-v2 — a deliberate new block, documented, never an ad-hoc edit mid-scene.
- Wardrobe is continuity. Within one story-day, the outfit is frozen in every prompt. The fastest way to break a sequence is describing the jacket in scene one and omitting it in scene three — omitted details don't persist; they re-roll.
- Pin the style too. A character block rendered "35mm photo" in one shot and "soft illustration" in the next reads as two characters. The style clause is part of the character.
- One character block per subject in multi-person scenes, maximally differentiated — distinct hair, distinct wardrobe colors, distinct accessories. Similar-looking characters in one frame invite feature swapping, and no block survives that.
- Test recognition, not beauty. The QA question for any new take isn't "does this look good?" but "would a viewer of scene one say this is the same person?" Squint-test thumbnails side by side.
Brands running a recurring mascot rather than a human character face the same mechanics with looser realism constraints — consistent character videos for brand mascots works that variant end to end.
FAQ
How long should a character description block be?
25–45 words, structured like a casting sheet: age range, skin tone, hair color and style, glasses/accessories, and specific wardrobe. Long enough to load the high-stickiness traits, short enough that it pastes into every scene prompt without crowding out action and camera language.
Why does my character's face change even with identical prompts?
Text cannot pin exact facial geometry — the model samples a plausible face matching your description each time, and identical words still allow a range of strangers. Build recognition from hair, wardrobe, and accessories instead, or escalate to image anchors and reference-to-video models, which pin the face by example rather than description.
Can I use the same character block across image and video models?
Yes, and you should — the block is model-agnostic casting information. Expect stickiness to vary by model, and always keep the style clause attached, because the same block rendered photoreal versus illustrated reads as two different characters.
When is text-only consistency good enough?
When appearances are spaced out (a character returning across separate posts), when the style is illustrated rather than photoreal, or when the character is incidental. For serialized content where the same person appears shot after shot, budget for reference-based generation from the start — it's cheaper than the re-roll lottery.
How do I handle two recurring characters in one scene?
Give each an atomic block with maximally different high-stickiness traits — different hair, different wardrobe palette, different accessories — and one action each. Visual similarity between characters is the main trigger for feature swapping, so differentiate harder than feels necessary.
Cast your character once, properly: write the block, generate the anchor portrait, and run your first reference-to-video scene on Versely — the character you save is the retake budget you keep.