AI model finder
Pick a model by what you can actually hand it. 148 models, 77 of them ranked, 9 that accept two or more reference images. Filter by input requirement rather than by name — the name is the least reliable signal there is.
148 of 148
| Model | Provider | Rank | Needs | Refs | Durations |
|---|---|---|---|---|---|
| GPT Image 2 Text to Image | OpenAI | #1 | prompt | — | — |
| Gemini 3.1 Flash TTS | #3 | prompt | — | — | |
| Nano Banana 2 | #3 | prompt | 14 | — | |
| Happy Horse 1.0 Text to Video | Alibaba | #3 | prompt | — | 3s, 4s, 5s, 6s |
| Seedance 2.0 | ByteDance | #3 | video ok | — | 4s, 5s, 6s, 7s |
| Mai Image 2.5 Edit | Microsoft | #3 | image | — | — |
| Cartesia Sonic 3.5 | Cartesia | #5 | prompt | — | — |
| Seedream 5.0 Pro | ByteDance | #6 | prompt | — | — |
| Wan 2.7 Text to Video | Wan | #6 | prompt | — | 2s, 3s, 4s, 5s |
| Happy Horse 1.1 Image to Video | Alibaba | #7 | image | — | 3s, 4s, 5s, 6s |
| Inworld TTS 1.5 Max | Inworld | #7 | prompt | — | — |
| Grok Imagine Image Quality | Grok | #9 | prompt | — | — |
| Grok Imagine Video | Grok | #9 | prompt | — | 6s, 10s, 15s, 20s |
| Inworld TTS 2 | Inworld | #9 | prompt | — | — |
| Kling 2.5 Turbo | Kling | #9 | prompt | — | 5s, 10s |
| HunyuanImage 3.0 Instruct Edit | Hunyuan | #10 | image | 3 | — |
| HiDream O1 Image | HiDream | #11 | prompt | — | — |
| ElevenLabs Multilingual | KIE | #11 | prompt | — | — |
| Luma UNI 1 Max | Luma | #11 | prompt | — | — |
| Flux 2 Max | Flux | #12 | prompt | — | — |
| Vidu Q3 Image to Video | Vidu | #12 | image | — | 5s, 10s, 15s |
| Vidu Q3 Video | Vidu | #12 | prompt | — | 5s, 10s, 15s |
| Kling Image 3.0 | Kling | #13 | prompt | — | — |
| Pixverse 5.6 Image to Video | Pixverse | #14 | image | — | 5s, 8s, 10s |
| Runway Gen-4.5 | Runway | #14 | prompt | — | 5s, 10s |
| VEO 3.1 | #15 | prompt | — | 4s, 6s, 8s | |
| Seedream 4.0 | ByteDance | #16 | prompt | — | — |
| Pixverse 5.6 Text to Video | Pixverse | #16 | prompt | — | 5s, 8s, 10s |
| Kling Image O1 | Kling | #16 | image | 10 | — |
| Wan 2.7 Image to Video | Wan | #17 | image | — | 2s, 3s, 4s, 5s |
| Wan 2.7 Pro Edit | Wan | #17 | image | — | — |
| Pixverse V6 Text to Video | Pixverse | #18 | prompt | — | 5s, 8s |
| Kling Video V3 Standard Image to Video | Kling | #18 | image | — | 3s, 4s, 5s, 6s |
| HiDream O1 Image Edit | HiDream | #19 | image | — | — |
| Flux 2 Flex | Flux | #19 | prompt | — | — |
| Pixverse V4 | Pixverse | #20 | prompt | — | 5s, 8s |
| Kling Video V2.6 Pro Text to Video | Kling | #24 | prompt | — | 5s, 10s |
| Ideogram V4 | Ideogram | #24 | prompt | — | — |
| Kling O3 Pro Text to Video | Kling | #25 | prompt | — | 3s, 4s, 5s, 6s |
| Krea 2 Medium | Krea | #26 | prompt | — | — |
| Wan V2.6 Text to Video | Wan | #28 | prompt | — | 5s, 10s, 15s |
| Wan 2.6 Text to Image | Wan | #29 | prompt | — | — |
| Luma Ray 2 720p | Luma | #29 | prompt | — | 5s, 9s |
| Flux 2 Klein 9B Base | Flux | #29 | prompt | 4 | — |
| Hailuo 2.3 Pro | Hailuo | #29 | image | — | 6s, 10s |
| Hailuo 2.3 Fast | Hailuo | #31 | image | — | 6s, 10s |
| Qwen Image Edit 2511 | Qwen | #31 | image | 6 | — |
| HiDream O1 Image Dev | HiDream | #32 | prompt | — | — |
| Seedance 1.5 Pro Text to Video | ByteDance | #33 | prompt | — | 4s, 5s, 6s, 7s |
| Seedance Image to Video | ByteDance | #34 | image | — | 2s, 3s, 4s, 5s |
| ERNIE Image Turbo | Baidu | #34 | prompt | — | — |
| Pixverse Text to Video | Pixverse | #35 | prompt | — | 5s, 8s, 10s |
| Recraft 4.1 Text to Image | Recraft | #37 | prompt | — | — |
| ImagineArt 2.0 | Imagine | #37 | prompt | — | — |
| Wan 2.6 Image to Video | Wan | #39 | image | — | 5s, 10s, 15s |
| Kling V2.1 | Kling | #40 | image | — | 5s, 10s |
| Flux 2 Flash | Flux | #41 | prompt | — | — |
| LTX 2.3 Text to Video Fast | LTX | #41 | prompt | — | 6s, 8s, 10s, 12s |
| LTX 2 Text to Video Pro | LTX | #42 | prompt | — | 6s, 8s, 10s |
| LTX 2 Pro | LTX | #43 | image | — | 6s, 8s, 10s |
| LTX 2 Text to Video Fast | LTX | #43 | prompt | — | 6s, 8s, 10s, 12s |
| Recraft V4 | Recraft | #45 | prompt | — | — |
| LTX 2 | LTX | #45 | prompt | — | 6s, 8s, 10s, 12s |
| Flux 2 Klein 4B Base Edit | Flux | #49 | image | 4 | — |
| Midjourney V7 Image to Video | Midjourney | #49 | image | — | — |
| PrunaAI P-Video | PrunaAI | #51 | prompt | — | 1s, 2s, 3s, 4s |
| LTX 2.3 Image to Video Pro | LTX | #51 | image | — | 6s, 8s, 10s |
| Flux Kontext | Flux | #56 | image | 1 | — |
| Runway Text to Video | Runway | #66 | prompt | — | 5s, 10s |
| Imagen 4 | #70 | prompt | — | — | |
| Qwen Image | Qwen | #79 | prompt | 1 | — |
| Qwen 3 TTS 0.6B | Qwen | #80 | prompt | — | — |
| Flux 1.1 Pro | Flux | #82 | prompt | — | — |
| Midjourney V7 | Midjourney | #84 | prompt | 5 | — |
| Qwen Z Image | Qwen | #103 | prompt | — | — |
| Flux Schnell | Flux | #121 | prompt | — | — |
| Runway Gen4 Image | Runway | #122 | image | — | — |
| Flux 3 Text to Video | Black Forest Labs | — | prompt | — | 5s, 6s, 7s, 8s |
| HeyGen Avatar V3 | HeyGen | — | prompt | — | — |
| HeyGen Avatar V5 | HeyGen | — | prompt | — | — |
Showing the first 80 of 148. Narrow the filters to see the rest.
Rank is the model's position in Versely's overall buyer ranking; most of the catalog is unranked, which is not a mark against it — ranking only covers models compared head to head. A model without its own page is folded into a parent family in the catalog.
Related
- Credit calculator — what a job on the model you picked will cost.
- Prompt builder — the parameters that model actually accepts, and how it likes to be prompted.
- Head-to-head comparisons — once the shortlist is down to two.
- The full catalog
FAQ
- Why filter by input instead of by name?
- Because the name almost never tells you what a model needs. A slug that sounds like an editing model can be filed as text-to-image only, with no reference support at all, and shopping by name is what produces a rejected job. Input requirements are a checkable field — whether it requires an image, accepts video, and how many reference images it takes — so that is what the filters use.
- What does the rank column mean?
- It is the model's position in Versely's overall buyer ranking. Most of the catalog is unranked, and that is not a mark against a model — ranking only covers models that have been compared head to head. An unranked model can be the right pick for a specific job.
- Why do some models have no page to click through to?
- The catalog folds variants into a parent family, so a model like an edit-mode or turbo variant is documented on its family's page rather than its own. Models whose catalog record is too thin to say anything useful also do not get a page. Both are deliberate.
- What is a reference image budget?
- The number of reference images a model will accept in one call. It matters when you need to lock more than one thing at once — a product, a face and a set, for example. A model with a budget of one cannot hold three references steady no matter how the prompt is written.