Happy Horse 1.0 Reference to Video: the stack is the brief (3s, 4s, 5s, 1080p, 28cr)
Happy Horse 1.0 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far.
Happy Horse 1.0 Reference to Video is reference-to-video. Upload the stills that must appear. Do not spend the whole budget because the slider goes that far. The catalog description is specific: it “generates videos from up to 9 reference images using character1–character9 placeholders in the prompt, with native synchronized audio and multilingual lip-sync.” That stack is the brief. The paragraph is only how the stack moves.
Happy Horse 1.0 Reference to Video sits on the Happy Horse roster (Alibaba parent). It requires an image. Max output is 1080p, with 720p also on the quality list. Aspect ratios are 16:9, 9:16, 1:1, 4:3, and 3:4. The credit sticker is 28. Durations run 3s through 15s in one-second steps. The slider going to 15s is not a reason to buy 15s.
character1 through character9 are the prompt
Nine stills is the ceiling, not a starter pack. The placeholders in the prompt are how you point at those stills. “character1 walks left, character2 holds the bottle” only works if you uploaded those two stills and they agree with themselves. A mismatched wardrobe across three “character1” frames is not extra coverage. It is an argument the model will split.
Reusable characters and products is the habit that makes a nine-slot stack usable more than once. Upload the cast once. Freeze it. New scenes reuse the same keys. If you rebuild the stack every generate, you will get nine cousins instead of one cast.
This is not Happy Horse 1.0 text-to-video. That sibling invents from words. This row is for the job where the face, the SKU, or the room must be the one in the stills. If you do not have the stills, you do not have this job. Make them, lock them, then open the reference row.
3s is a continuity check. 15s is a spend.
The duration list is 3s, 4s, 5s, 6s, 7s, 8s, 9s, 10s, 11s, 12s, 13s, 14s, 15s. Start at 3s, 4s, or 5s while you prove the stack holds: same face, same product, same wardrobe. A 15-second take of a stack that already drifted at second two is 13 more seconds of the wrong person.
1080p is the ceiling on this row. There is no 4K token here. If the delivery spec is 4K, this is the wrong generate — or you upscale after picture lock, which is a different model. For most short-form, 1080p is the ship resolution. Do not treat 720p as a failure; it is on the quality list for a reason.
The 28-credit sticker is per generate, not per still in the stack. Filling all nine slots does not make the clip more “worth” 15 seconds. It makes the identity contract denser. Identity is usually proven in five seconds. Action that needs fifteen should be a scene you already trust at five.
Native audio and lip-sync are in the description. Use them or they use you.
The catalog does not flip a separate audio boolean on this record, but the description names native synchronized audio and multilingual lip-sync. A mute prompt on a talking stack still tends to come back with a mouth and a bed. Write the line. Name the language. If the shot is silent product motion, say no speech, no song — or you will spend 28 credits on a presenter you will strip.
Multilingual lip-sync is not “translate later in captions.” It is a generate-time mouth. If the campaign needs the same face in two languages, that is two generates against the same frozen stack, two scripts, not one clip with burned-in subtitles pretending to be a dub.
The category index is best reference-to-video. The generate door for a scene that is not reference-locked is the AI video generator. This page is only the Happy Horse 1.0 reference row: up to nine stills, character1–character9, 3s–15s, 1080p, 28 credits, Alibaba. Upload the cast. Then move it. Do not rent fifteen seconds because the slider offered them.
FAQ
Do I have to use all nine reference images?
No. Nine is the maximum. Three clean, agreeing stills beat nine that argue. Use extra slots for a second character or a product that must appear, not for more angles of a face you already locked.
Does this row require an image?
Yes. The catalog marks it requires_image. A text-only prompt belongs on the text-to-video sibling, not here.
Why not always generate 15 seconds at 1080p?
Because 15s is the ceiling and 1080p is the ceiling, not a package deal you owe every take. Prove identity at 3s–5s. Lengthen when the action, not the anxiety, needs it.
Is the audio a separate model?
Not on this row. The description includes native synchronized audio and multilingual lip-sync as part of the generate. Write the sound into the prompt. A later TTS pass will not match a mouth this row already animated.