Guides

    How to hold a product in hand without inventing a creator

    Photograph the real hand and the real SKU, then I2V on Seedance 2.5 via `/tools/ai-image-to-video` (4–30s, 720p, 47cr, native audio). A generated creator holding a cousin bottle is a different identity you do not own.

    Versely Team6 min read

    Photograph the real hand and the real SKU, then I2V on Seedance 2.5 via /tools/ai-image-to-video (4–30s, 720p, 47cr, native audio). A generated creator holding a cousin bottle is a different identity you do not own.

    The UGC-shaped product ad fails in two places at once. The face is a person you did not hire and cannot clear. The pack is a near-miss: cap colour off, type redrawn, a seam the real bottle does not have. Text-to-video will invent both in one pass if you let it. The fix is not a better "pretty woman holds serum" prompt. The fix is a photograph of a real hold, then motion on a row that keeps frame one.

    Why the generated creator is not the job

    UGC originally meant a customer or fan made the file. The feed now treats "handheld, kitchen light, product in frame" as a look. A generated presenter is a likeness you do not have a release for. Even when the face is generic, you are shipping a person who does not exist as if they used the product.

    The bottle is the second identity. If the model restages the scene, it will also restage the pack. You asked for a hold. You received a cousin SKU in a cousin kitchen, performed by a cousin hand. That file is unusable as product creative no matter how good the motion is.

    You do not need a creator. You need a hand, a pack, and a first frame. POV hands — wrists in, face out — are enough for most holds. A staff hand with a release is enough when you want more body in the shot. A hired creator is a different budget. Generate is not a substitute for that hire.

    The product video generator is the tool when the brief is "this object, new scene." This page is narrower: this object, this hand, motion only.

    Shoot the still as if it were the ad

    Light the pack so the label reads. Hold it the way a person actually holds it — not a claw, not a floating grip, not eight fingers. Fill the frame enough that the wordmark survives a 9:16 crop. Phone camera is fine. A dirty kitchen is fine. A seamless sweep is fine if that is the brand. What is not fine is a prompt that describes the kitchen instead of a JPEG of it.

    Checklist before you leave the table:

    • The SKU in the photo is the SKU on the shelf. Same colourway, same cap, same type.
    • The hand is a real hand you can account for: yours, a teammate's, a creator you paid.
    • The label is sharp. Motion will not invent missing type.
    • The first frame could be the poster. If you would not post the still, do not buy seconds.

    Shoot two or three angles: three-quarter, label-on, cap twist starting. Pick one still per generate. Image-to-video treats that still as frame one. If you want a new composition, shoot a new still. Do not ask the video row to recast the hold.

    Skip the face unless you have a person and a reason. A forehead and a smile the model invented are how a product hold becomes an unauthorized presenter. Hands plus pack is the whole format.

    Animate on Seedance 2.5, not a restage

    Seedance 2.5 is the published row: 4–30 seconds in one pass, 720p ceiling, listed 47 credits, native audio. The catalog category is text-to-video. The still-to-motion job on that family runs through image-to-video. Do not invent /models/seedance-2-5-image-to-video. That slug is not a live page. The parent model stays /models/seedance-2-5.

    Prompt only what should move:

    Slow turn of the bottle in the same hand. Fingers adjust on the cap. Keep the label readable. Same kitchen, same light. No new person. No new pack.

    Do not prompt a review, a cafe the still does not show, or "UGC creator talks about ingredients." That sentence invents a mouth on a hold you already photographed. If you need a line, caption it or record a voice you own.

    Duration: pick the hold you can actually perform. A cap twist is four to eight seconds. Thirty seconds is the ceiling, not a requirement. 720p is the ceiling on this row. Native audio is the mix Seedance writes with the picture — room, pack rustle, a cap click — not a spokesperson you did not cast. If the mix is wrong, mute it. Do not re-roll identity to chase a whoosh.

    If the output restages the bottle, the still was too weak or the prompt asked for a new scene. Tighten the still. Shorten the motion sentence. Do not switch to text-to-video "to get more energy." Energy that costs you the SKU is not energy you can ship.

    What you should not do instead

    Do not text-to-video the whole ad. A paragraph that names the brand and a "real person" will give you a person and a pack the legal team cannot sign. That is the cousin-bottle failure with extra charm.

    Do not generate a creator, then inpaint your product into their hand. You still do not own the face. You now have a compositing job on top of a likeness problem.

    Do not use reference-to-video when the frame is already the deliverable. Reference-to-video recasts. Image-to-video continues. A hold shot is a continuation job.

    Do not hire nobody and imply a user. A generated peer review is not UGC. It is an ad wearing a peer's clothes. If you need a real recommendation, ask a customer or pay a creator. If you need a pack in motion, photograph the hold.

    The workflow is three files: a JPEG you would post, a Seedance take that only moves, a caption or a voice you wrote. None of those files require a fictional influencer. The hand is already in the room.

    FAQ

    Can I generate the hand if I do not want my face in the ad?

    You can omit the face. You should not invent the hand and the pack. Photograph a real hold with the real SKU, crop to wrists if you want, then image-to-video. A generated hand holding a generated bottle is two identities you do not own.

    Why Seedance 2.5 instead of a short text-to-video clip?

    Because the still is the contract. Seedance 2.5 is 4–30s, 720p, listed 47 credits, native audio, I2V from the JPEG you approved. Text-to-video invents the hold. That is how the cousin bottle arrives.

    Does native audio mean I get a voiceover for free?

    No. Native audio on this row is the mix generated with the picture — ambience and foley, not a hired host. If you need a line, write it, record it, or caption it. Do not prompt a talking creator into a still that has no mouth.

    Is a product-in-hand generate the same as UGC?

    No. UGC is content a real user made, or paid creator work built to look like it. A photographed hold you animate is product footage. It can look handheld. It is not a user recommendation unless a user actually made it.