Models and architecture

    Elo rating

    Also called Elo score, Arena ranking.

    An Elo rating ranks models by head-to-head preference: people compare two outputs from the same prompt, and each model's number moves according to who won and how strong the opponent was.

    It comes from chess, and the borrowed logic is the point. Beating a highly-rated opponent moves your number more than beating a weak one, so a rating reflects the quality of the field a model has faced rather than a raw win count. It is a relative measure — it only means something against the population it was computed over.

    It exists because the alternatives are worse. Automated image metrics correlate poorly with what people actually prefer, and a single expert's opinion does not generalise. Aggregated blind pairwise votes are noisy per vote and reasonably stable in bulk.

    The caveats are real. Ratings drift as new models enter, prompt mix changes what is being measured, and preference rewards what looks striking in a side-by-side — which is not always what survives in a finished edit. A ranking is a starting shortlist, not a verdict on your brief.

    In practice

    • Relative, not absolute: a number only means something against the same field.
    • Small gaps between neighbouring models are usually inside the noise.
    • Category rankings beat overall rankings when your job is specific.

    Models carrying a ranking

    Catalog entries with at least one placed result in the live rankings. 148 of the 296 models in the Versely catalog qualify.

    ModelProviderType
    GPT Image 2 Text to ImageOpenAIImage
    Mai Image 2.5 EditMicrosoftImage
    Happy Horse 1.0 Text to VideoAlibabaVideo
    Gemini 3.1 Flash TTSGoogleAudio
    Nano Banana 2GoogleImage
    Seedance 2.0ByteDanceVideo
    Cartesia Sonic 3.5CartesiaAudio
    Seedream 5.0 ProByteDanceImage

    Browse all 77 spec pages for full settings, resolutions and credit costs.

    The mistake to avoid

    Picking the top-rated model for every job. Ratings aggregate across prompt types, and the leader on general preference is often not the leader on your particular one.

    Go deeper

    ELO Rankings for AI Models, Explained

    How ELO rankings for AI models work: pairwise voting, rating math, what leaderboard gaps mean, and the blind spots to know before trusting a score.

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.