Wan 2.1 Image to Video: the still is the contract (5s, 10s, 15s, 1080p, 80cr)
Wan 2.1 Image to Video is image-to-video. If the label is wrong, every second is waste.
Wan 2.1 Image to Video is image-to-video. If the label is wrong, every second is waste.
Wan 2.1 Image to Video is an open-source AI video generation model that utilizes a diffusion transformer architecture and a novel 3D spatio-temporal VAE (Wan-VAE) for image-to-video generation. The category is image-to-video. requires_image is true. Audio is empty — this is a silent row. Durations: 5s, 10s, 15s. Max output: 1080p. Catalog credits: 80, and the row is marked discounted. Architecture trivia does not replace a plate. The still is the contract.
The upload is not optional
requires_image: true means the generate will not start from a paragraph. Frame one is the file. If that file is a muddy screenshot of the product, you will get 5, 10, or 15 seconds of muddy product. Wan-VAE will not typeset a label that was never in the pixels.
Lock the still on a stills row first. Then open image-to-video and describe only motion: orbit, steam, a slow push, a head turn. Re-describing the object is how i2v models morph the subject they can already see.
Alibaba's video family roster is Wan. Best Wan model is the ranking. This slug is the 2.1 image-to-video line, not a later Wan text-to-video sibling.
Silent 5, 10, or 15 at 1080p
There is no native audio on this row. Do not write dialogue into the prompt and expect a voice. If you need speech, plan a voiceover or a lipsync pass after, or pick a video model whose audio field is true.
Duration is a three-way choice: 5s, 10s, 15s. Fifteen seconds of a bad plate is three times the waste of five. Start at 5s until the identity holds. Qualities listed: 480p, 720p, 1080p. The ceiling is 1080p. Do not prompt 4K.
Eighty credits is the catalog figure. The row is discounted in the snapshot; the live quote in the app is the number that will actually charge. Do not convert either figure to USD. Do notice that this is not a 2-credit sketch. You are paying for motion of a picture you already had.
Diffusion transformer is not a prompt-only row
Wan-VAE and a diffusion transformer are why the motion can look coherent. They are not why you can skip the still. Open-source architecture is a provenance note, not a workflow. The workflow is:
- Approve the still.
- Upload it.
- Pick 5s / 10s / 15s and 1080p (or a lower quality if that is the job).
- Prompt camera and subject motion only.
If step 1 did not happen, stop. Best image-to-video model is the category ranking if you are still choosing a family. Wan 2.1 i2v is the family member that demands the file.
Discounted is not throwaway
Discounted credits still buy seconds of the wrong product. A cheap 15-second morph is not a bargain. It is a 15-second morph. Hold the still at full size. If you would not ship it as a thumbnail, do not spend 80 credits teaching it to move.
Wan-VAE will move whatever you handed it, including the typo, for 5, 10, or 15 silent seconds at 1080p.
FAQ
Can I run Wan 2.1 Image to Video from text only?
No. requires_image is true and the category is image-to-video. Without a plate, pick a text-to-video Wan row instead of hoping this one will invent the product.
Does it generate sound?
The catalog audio field is empty. Treat the output as silent. Add voice, music, or lipsync on a different row if the brief needs a track.
What lengths can I generate?
5 seconds, 10 seconds, or 15 seconds, at up to 1080p. There is no 3-second or 30-second option on this slug.
Why is the credit number 80 if the row is discounted?
80 is the catalog credits value on the model. The snapshot also marks the row discounted. Use the in-app quote for the charge. Neither number is a USD price, and neither number is a reason to skip the still.