Generation modes

    Reference-to-video

    Also called R2V.

    Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.

    The distinction from image-to-video matters. An image-to-video input is the first frame — its framing, angle and lighting are locked in. A reference is not a frame at all: it is an identity the model may place anywhere in the shot, at any distance, from any angle, under whatever light the prompt asks for.

    That is what makes it the mode for recurring subjects. Hand it a product from three angles, or a character and the set they live in, and the same thing can appear in shot after shot without you re-shooting the opening frame each time.

    The ceiling is how many references a model will hold at once. That number is published per model and is a hard limit — extra images are ignored, not blended, so a five-reference brief on a two-reference model quietly drops three of them.

    In practice

    • References work best when they disagree usefully: different angles beat three near-identical crops.
    • The prompt still owns the shot — references decide who is in it, not what happens.
    • Backgrounds inside a reference image leak. Clean plates and cutouts carry less of what you did not intend.

    Reference-to-video models

    Catalog entries that carry a subject from reference stills into motion. 17 of the 296 models in the Versely catalog qualify.

    The mistake to avoid

    Submitting more references than the model accepts and assuming they were all used. Check the model's reference limit before blaming the output.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.