Versely

    AI model finder

    Pick a model by what you can actually hand it. 148 models, 77 of them ranked, 9 that accept two or more reference images. Filter by input requirement rather than by name — the name is the least reliable signal there is.

    148 of 148
    ModelProviderRankNeedsRefsDurations
    GPT Image 2 Text to ImageOpenAI#1prompt
    Gemini 3.1 Flash TTSGoogle#3prompt
    Nano Banana 2Google#3prompt14
    Happy Horse 1.0 Text to VideoAlibaba#3prompt3s, 4s, 5s, 6s
    Seedance 2.0ByteDance#3video ok4s, 5s, 6s, 7s
    Mai Image 2.5 EditMicrosoft#3image
    Cartesia Sonic 3.5Cartesia#5prompt
    Seedream 5.0 ProByteDance#6prompt
    Wan 2.7 Text to VideoWan#6prompt2s, 3s, 4s, 5s
    Happy Horse 1.1 Image to VideoAlibaba#7image3s, 4s, 5s, 6s
    Inworld TTS 1.5 MaxInworld#7prompt
    Grok Imagine Image QualityGrok#9prompt
    Grok Imagine VideoGrok#9prompt6s, 10s, 15s, 20s
    Inworld TTS 2Inworld#9prompt
    Kling 2.5 TurboKling#9prompt5s, 10s
    HunyuanImage 3.0 Instruct EditHunyuan#10image3
    HiDream O1 ImageHiDream#11prompt
    ElevenLabs MultilingualKIE#11prompt
    Luma UNI 1 MaxLuma#11prompt
    Flux 2 MaxFlux#12prompt
    Vidu Q3 Image to VideoVidu#12image5s, 10s, 15s
    Vidu Q3 VideoVidu#12prompt5s, 10s, 15s
    Kling Image 3.0Kling#13prompt
    Pixverse 5.6 Image to VideoPixverse#14image5s, 8s, 10s
    Runway Gen-4.5Runway#14prompt5s, 10s
    VEO 3.1Google#15prompt4s, 6s, 8s
    Seedream 4.0ByteDance#16prompt
    Pixverse 5.6 Text to VideoPixverse#16prompt5s, 8s, 10s
    Kling Image O1Kling#16image10
    Wan 2.7 Image to VideoWan#17image2s, 3s, 4s, 5s
    Wan 2.7 Pro EditWan#17image
    Pixverse V6 Text to VideoPixverse#18prompt5s, 8s
    Kling Video V3 Standard Image to VideoKling#18image3s, 4s, 5s, 6s
    HiDream O1 Image EditHiDream#19image
    Flux 2 FlexFlux#19prompt
    Pixverse V4Pixverse#20prompt5s, 8s
    Kling Video V2.6 Pro Text to VideoKling#24prompt5s, 10s
    Ideogram V4Ideogram#24prompt
    Kling O3 Pro Text to VideoKling#25prompt3s, 4s, 5s, 6s
    Krea 2 MediumKrea#26prompt
    Wan V2.6 Text to VideoWan#28prompt5s, 10s, 15s
    Wan 2.6 Text to ImageWan#29prompt
    Luma Ray 2 720pLuma#29prompt5s, 9s
    Flux 2 Klein 9B BaseFlux#29prompt4
    Hailuo 2.3 ProHailuo#29image6s, 10s
    Hailuo 2.3 FastHailuo#31image6s, 10s
    Qwen Image Edit 2511Qwen#31image6
    HiDream O1 Image DevHiDream#32prompt
    Seedance 1.5 Pro Text to VideoByteDance#33prompt4s, 5s, 6s, 7s
    Seedance Image to VideoByteDance#34image2s, 3s, 4s, 5s
    ERNIE Image TurboBaidu#34prompt
    Pixverse Text to VideoPixverse#35prompt5s, 8s, 10s
    Recraft 4.1 Text to ImageRecraft#37prompt
    ImagineArt 2.0Imagine#37prompt
    Wan 2.6 Image to VideoWan#39image5s, 10s, 15s
    Kling V2.1Kling#40image5s, 10s
    Flux 2 FlashFlux#41prompt
    LTX 2.3 Text to Video FastLTX#41prompt6s, 8s, 10s, 12s
    LTX 2 Text to Video ProLTX#42prompt6s, 8s, 10s
    LTX 2 ProLTX#43image6s, 8s, 10s
    LTX 2 Text to Video FastLTX#43prompt6s, 8s, 10s, 12s
    Recraft V4Recraft#45prompt
    LTX 2LTX#45prompt6s, 8s, 10s, 12s
    Flux 2 Klein 4B Base EditFlux#49image4
    Midjourney V7 Image to VideoMidjourney#49image
    PrunaAI P-VideoPrunaAI#51prompt1s, 2s, 3s, 4s
    LTX 2.3 Image to Video ProLTX#51image6s, 8s, 10s
    Flux KontextFlux#56image1
    Runway Text to VideoRunway#66prompt5s, 10s
    Imagen 4Google#70prompt
    Qwen ImageQwen#79prompt1
    Qwen 3 TTS 0.6BQwen#80prompt
    Flux 1.1 ProFlux#82prompt
    Midjourney V7Midjourney#84prompt5
    Qwen Z ImageQwen#103prompt
    Flux SchnellFlux#121prompt
    Runway Gen4 ImageRunway#122image
    Flux 3 Text to VideoBlack Forest Labsprompt5s, 6s, 7s, 8s
    HeyGen Avatar V3HeyGenprompt
    HeyGen Avatar V5HeyGenprompt

    Showing the first 80 of 148. Narrow the filters to see the rest.

    Rank is the model's position in Versely's overall buyer ranking; most of the catalog is unranked, which is not a mark against it — ranking only covers models compared head to head. A model without its own page is folded into a parent family in the catalog.

    Related

    FAQ

    Why filter by input instead of by name?
    Because the name almost never tells you what a model needs. A slug that sounds like an editing model can be filed as text-to-image only, with no reference support at all, and shopping by name is what produces a rejected job. Input requirements are a checkable field — whether it requires an image, accepts video, and how many reference images it takes — so that is what the filters use.
    What does the rank column mean?
    It is the model's position in Versely's overall buyer ranking. Most of the catalog is unranked, and that is not a mark against a model — ranking only covers models that have been compared head to head. An unranked model can be the right pick for a specific job.
    Why do some models have no page to click through to?
    The catalog folds variants into a parent family, so a model like an edit-mode or turbo variant is documented on its family's page rather than its own. Models whose catalog record is too thin to say anything useful also do not get a page. Both are deliberate.
    What is a reference image budget?
    The number of reference images a model will accept in one call. It matters when you need to lock more than one thing at once — a product, a face and a set, for example. A model with a budget of one cannot hold three references steady no matter how the prompt is written.