Guides

    Happy Horse 1.1 Image to Video: the still is the contract (3s, 4s, 5s, 1080p, 18cr)

    Happy Horse 1.1 Image to Video is image-to-video. If the label is wrong, every second is waste.

    Versely Team4 min read

    Happy Horse 1.1 Image to Video is image-to-video. If the label is wrong, every second is waste. Alibaba's row animates a first-frame image into 1080p video with synchronized native audio and multilingual lip-sync. Aspect ratio is inferred from the image. Duration is 3s through 15s. Catalog price is 18 credits. The still is frame one. The prompt is motion and speech, not a new set.

    This is not reference-to-video. You do not get a new beach behind the bottle. You get the bottle, in the room you photographed, starting to move. Pick the category on purpose. The image-to-video tool is the door. The Happy Horse roster is the family.

    Frame one is the only frame you lock

    Requires an image. No still, no job. Composition, wardrobe, label, and background are already decided before you press generate. If the hero is the wrong color, do not prompt "make the bottle blue." Fix the still, then animate. A text-to-video row would invent the bottle; this row is contracted to the pixels you sent.

    Crop for the platform before you upload. The model infers aspect from the file. A 16:9 still does not become a 9:16 Reel because you asked. Export the still at the ratio you will ship.

    Quality on the row is 720p and 1080p. Max output is 1080p. Do not write "4K" into the prompt and expect a setting the catalog does not list.

    Speech is generated with the picture

    The audio field on the snapshot is empty, but the description is not coy: synchronized native audio and multilingual lip-sync. Treat that as a soundtrack you will receive. A talking still with no line in the prompt still gets a voice you did not cast. Write the language and the line. If the mouth must stay shut, say that the subject does not speak.

    This is why the row exists next to dedicated lipsync models. Happy Horse creates a speaker from a still. It is not how you retrofit an already-shot clip. For that job see Sync Lipsync 2.0 or the lipsync tool.

    3s is a beat; 15s is the ceiling

    Durations run 3s, 4s, 5s, then every second up to 15s. There is no 2s option here — that is a different Wan row. There is no 20s option. Start at 3s when you are testing the still. Pay 18 credits for a 15s take only after the 3s version already has the right face and the right line.

    The best image-to-video ranking is the rest of the category. Image-to-video vs text-to-video is the decision if you do not yet own a frame. This page is only Happy Horse 1.1: first-frame contract, 1080p, 3–15s, 18 credits, native audio in the description, multilingual mouth.

    If the still is wrong, every extra second is a more expensive wrong clip. Lock the picture. Then pick a length.

    FAQ

    Can I restage the product in a new location?

    Not on this row. Image-to-video keeps the uploaded frame as the opening picture. For a new scene with the same SKU, use a reference-to-video model and a clean product still.

    Does Happy Horse 1.1 require an image?

    Yes. requires_image is true. Text-only prompts belong on a text-to-video slug.

    What lengths can I actually pick?

    3s through 15s, in one-second steps. The short title lists 3s, 4s, 5s because those are the cheap tests. 15s is available. 1s and 2s are not.

    Is the 18-credit price a silent clip?

    No. The catalog description includes synchronized native audio and multilingual lip-sync. Prompt the language and whether anyone talks. A silent brief still tends to come back with sound you did not choose.