Workflows

    Mascot Sheets: Turnarounds, Expressions, and Pose Libraries

    Produce the reference sheet as its own deliverable first, then feed it into every downstream generation — the model-sheet habit animation never dropped.

    Versely Team8 min read

    Comic-style character panels laid out together showing consistent poses and expressions

    Traditional character design never generates a character by describing it once per shot. Before a single frame of production art gets made, there's a model sheet: the character from the front, from a three-quarter angle, from the side and back, plus a grid of expressions and a set of signature poses — one reference document, built once, that every artist who touches the character afterward works from. AI generation makes it tempting to skip that step, because a prompt is so cheap to write that redescribing the mascot for every new shot feels faster than building a sheet first. It isn't, past about the third generation, because the redescription is exactly where drift creeps in. The old craft had the order right. Build the sheet. Then generate from it.

    Three sheets, three different jobs

    "Mascot sheet" isn't one artifact — it's three, and they solve different downstream problems:

    A turnaround is the character from multiple fixed angles — front, three-quarter, profile, back — in one consistent pose and outfit. Its job is giving every future generation a real 3D sense of the character's proportions and silhouette, so a shot from an angle you never explicitly generated before still looks like the same mascot rather than a new interpretation of one.

    An expression sheet is one pose, one angle, held constant, with the face cycled through a range of emotions — neutral, happy, surprised, concerned, annoyed is a reasonable minimum set for anything with dialogue or reaction beats. Its job is giving a video or dialogue-heavy campaign somewhere to pull a matching emotional register from, instead of re-describing "excited" from scratch every time and getting a slightly different face back.

    A pose library holds the face and outfit constant and varies the body — standing, walking, gesturing, mid-action. Its job is coverage for anything that needs the mascot doing something rather than just facing camera: an explainer video, a product demo, a comic panel.

    Most mascot campaigns need at least the turnaround and one of the other two. Which second sheet matters more depends entirely on the format — expression coverage for a talking mascot, pose coverage for one that's mostly shown doing things.

    Why the sheet has to exist before the campaign, not during it

    Nothing about a character carries between separate generations by default — every run starts fresh, and a face or a costume described in words gets quietly re-invented each time, close enough to feel deliberate and different enough that it reads as a different mascot three shots in. The fix that actually holds is supplying reference images instead of adjectives, and a sheet is the most efficient way to produce a full set of good references at once rather than one at a time.

    The part worth being deliberate about: generate the whole sheet as a single image — one grid containing all the angles or expressions in one generation — rather than as several separate generations you stitch together afterward. Consistency within one generation is dramatically stronger than consistency across several, because the model is holding one coherent scene in mind rather than reconstructing the character from a text description each time. A turnaround generated as four separate prompts for "front," "side," "three-quarter," "back" will drift between them in exactly the ways a single four-panel grid generation won't, because the second approach never actually held all four angles in the same frame at once.

    This matters more for mascots specifically than it does for photorealistic human spokespeople. A stylized, illustrated character has fewer real-world reference photos a model can lean on to fill gaps, which means the sheet is carrying more of the identity than it would for a photoreal face — there's no "looks roughly like a person" fallback to catch the model if the sheet is thin.

    Building each sheet in one pass

    The prompting pattern is the same shape for all three, and it's worth stating explicitly because it's easy to under-specify:

    • Name the layout, not just the content. "A four-panel turnaround sheet, front view, three-quarter view, side view, back view, arranged left to right, consistent lighting and proportions across all four panels" gives the model a structure to hold, rather than leaving it to infer that you wanted a grid at all.
    • Freeze everything except the one variable you're actually cycling. For an expression sheet, the pose, angle, outfit, and lighting should be identical across every panel — only the face changes. For a pose library, the face and outfit hold, only the body changes. Any variable you didn't explicitly freeze is a variable the model is free to drift on panel to panel.
    • Give expressions and poses as a specific named list, not a vague instruction to "show some emotions" or "a few different poses" — "neutral, happy, surprised, concerned, determined" produces a usable grid; "various expressions" produces five random guesses at what various means.
    • Generate at a resolution that survives cropping. A sheet only pays off once individual panels get cropped out and used as standalone references later, so it needs enough native resolution that a single panel still holds up on its own.

    Turning the sheet into references models can actually use

    Once the sheet exists, the individual panels — cropped out, or in some cases the whole grid — become the reference set that later generations point back to. How many of them a model can hold onto at once is a real, checkable ceiling, not a vague capacity claim: FLUX.2 maintains character consistency across up to 10 simultaneous reference images, and Kling Image O1 supports up to 10 reference images for feature consistency, alongside series content creation — headroom for a full turnaround plus a couple of expression panels in one generation call, rather than being limited to a single reference and hoping the rest gets improvised correctly.

    That headroom is also what makes a pose library specifically worth building for video work, not just images. Reference-to-video is its own live model category — 17 models deep in Versely's current catalog — built specifically to take a character's reference images and carry that identity into motion. A pose library gives that category something concrete to anchor to beyond a single static portrait: a body doing the kind of thing the video actually needs it to do, rather than a face the model has to guess a full body from.

    A Versely walkthrough: sheet first, then register it

    The practical sequence keeps the sheet-building step and the reuse step separate, because they're solving different problems — one produces good references, the other makes sure the rest of a production keeps pointing back at them.

    1. Generate the sheet as one image. "Generate a four-panel character turnaround sheet for our mascot — front, three-quarter, side, and back view, same lighting and proportions across all four panels" — one generation, one grid, maximum internal consistency. Repeat the same pattern for an expression sheet or pose library, depending on what the campaign actually needs coverage for.
    2. Crop the panels you'll reuse. A single panel, isolated, is a cleaner reference than the full grid for most downstream generations — crop the specific angle or expression a given shot actually calls for.
    3. Register the cropped set as a reusable asset, so the rest of the production points at it by name instead of re-uploading or redescribing it per scene. Ask the agent to set up a reference asset for the mascot — that opens an upload card where the sheet's cropped panels get attached once, and the asset stays available across every scene in the workflow from then on.
    4. Generate downstream shots from the registered asset, not from a fresh description. A scene skeleton that points at the registered mascot and adds only what's new about that shot holds identity far better than one that re-describes the character's appearance from memory each time.

    FAQ

    Should the whole sheet be one generation or several?

    One, where possible. Consistency within a single generation is significantly stronger than consistency across separate ones, because the model holds all the panels in the same frame at once rather than reconstructing the character from a text description on every separate call.

    How many expressions actually need to be on the sheet?

    Five is a reasonable practical floor for anything with dialogue or reactions — neutral, happy, surprised, concerned, and one negative emotion like annoyed or sad. Add more only for a campaign that specifically needs a wider emotional range; a sparse, well-executed set of five beats a crowded twelve where several panels barely differ.

    Does a mascot sheet work the same way for a photorealistic spokesperson?

    The turnaround and pose-library logic carries over directly. The stakes are a little different: a photoreal face has more of a real-world fallback a model can lean on if references are thin, where a stylized mascot's identity is carried almost entirely by the sheet itself, which is exactly why the sheet matters more, not less, for illustrated characters.

    What if the sheet itself has an inconsistent panel?

    Reject and regenerate the whole sheet rather than trying to patch one panel in isolation — a mismatched panel spliced in from a separate generation reintroduces the exact cross-generation drift the single-image sheet was built to avoid.

    Build the sheet once, register the panels you'll reuse, and let every shot after that point back at pixels instead of a redescribed memory of the mascot.