AI Models

    330 AI models, one subscription.

    Every video, image, lipsync, upscaling and audio model Versely ships — real credit pricing, real rankings, no marketing fluff. Tier and input-mode variants of the same model share a page, so the 188 pages below cover the whole catalogue without repeating themselves.

    Browse by spec

    The same catalogue sliced by what you are actually choosing on — clip length, output quality, how a model bills, and what it can do. Each page prices the whole roster for that setting.

    Video models

    86 pages

    Text-to-video, image-to-video and video-to-video generation.

    AlibabaFeatured

    Happy Horse 1.0 Text to Video

    Happy Horse 1.0 generates expressive videos from text prompts with native synchronized audio, Foley sound effects, and multilingual lip-sync — supporting up to 1080p output and durations from 3 to 15 seconds

    34 credits+3 variants
    ByteDanceFeatured

    Seedance 2.0

    Seedance 2.0 generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability

    44 credits+7 variants
    WanFeatured

    Wan 2.7 Text to Video

    Wan 2.7 generates high-fidelity videos from text prompts with strong motion consistency, optional custom audio input, and intelligent prompt rewriting

    16 credits#7 overall
    AlibabaFeatured

    Happy Horse 1.1 Image to Video

    Happy Horse 1.1 animates a first-frame image into 1080p video with synchronized native audio and multilingual lip-sync (aspect ratio inferred from the image, 3–15s)

    34 credits#8 overall
    RunwayFeatured

    Runway Gen-4.5

    Runway Gen-4.5 video generation with state-of-the-art quality, cinematic motion, and precise camera controls

    200 credits#16 overall
    WanFeatured

    Wan 2.7 Image to Video

    Wan 2.7 animates images with three modes: first-frame to video, first-and-last-frame interpolation, or video continuation, with optional driving audio

    16 credits#16 overall
    PrunaAIFeatured

    PrunaAI P-Video

    Fast video generation with built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-to-video in a single endpoint. Generates 5s 720p video in ~10 seconds.

    2 credits per second of output#55 overall
    Black Forest LabsFeatured

    Flux 3 Text to Video

    FLUX.3 is Black Forest Labs' frontier video model — generates video with native audio directly from a text prompt, at up to 1080p and 5–20 seconds across eight aspect ratios

    34 credits
    MiniMaxFeatured

    Minimax H3 Text to Video

    MiniMax H3 is a frontier video model — generates 2K video from a text prompt alone, in durations from 5 to 15 seconds across six aspect ratios

    104 credits+4 variants
    ByteDanceFeatured

    Seedance 2.5

    Dreamina Seedance 2.5 generates a native 30-second single-shot video at up to 720p from one text prompt, reasoning about the whole shot at once so motion, lighting and subject identity stay coherent from first frame to last.

    67 credits+2 variants
    WanFeatured

    Wan 3.0 Text to Video

    Alibaba's Wan 3.0 generates video with native audio from a text prompt. Ranked #1 on the Artificial Analysis text-to-video board at launch. 480p, 720p or 1080p, 2-15 seconds, five aspect ratios or adaptive, with prompt expansion and an optional slower reasoning pass.

    8 credits+2 variants
    Grok

    Grok Imagine Video

    Grok image-to-video generation

    24 credits#10 overall
    Kling

    Kling 2.5 Turbo

    Fast video generation with good quality balance

    17 credits#10 overall
    Vidu

    Vidu Q3 Image to Video

    Transform static images into dynamic videos with Vidu Q3 technology

    14 credits (discounted)+2 variants
    Pixverse

    Pixverse 5.6 Image to Video

    Latest Pixverse model for converting images to high-quality animated videos

    28 credits#15 overall
    Google

    VEO 3.1

    Enhanced VEO 3 with improved quality and features

    64 credits+3 variants
    Kling

    Kling Video V3 Standard Image to Video

    Kling V3 Standard image-to-video generation with camera controls and element composition

    51 credits+3 variants
    Vidu

    Vidu Q3 Video

    Generate high-quality videos from text prompts with Vidu Q3 engine

    28 credits#18 overall
    Pixverse

    Pixverse 5.6 Text to Video

    Generate videos directly from text prompts with Pixverse 5.6 advanced capabilities

    28 credits#19 overall
    Pixverse

    Pixverse V6 Text to Video

    PixVerse V6 text-to-video. Generates video from a text prompt with strong motion and prompt control, optional audio and multi-clip, up to 1080p.

    10 credits+1 variants
    Pixverse

    Pixverse V4

    Pixverse version 4 for creative video generation

    12 credits#21 overall
    Kling

    Kling Video V2.6 Pro Text to Video

    Professional text-to-video with Kling V2.6 Pro

    56 credits+3 variants
    Kling

    Kling O3 Pro Text to Video

    Kling O3 Pro text-to-video with reasoning-enhanced generation, camera controls, and audio

    56 credits+6 variants
    Hailuo

    Hailuo 2.3 Pro

    Professional-grade video generation with enhanced features

    40 credits#30 overall
    Hailuo

    Hailuo 2.3 Fast

    Fast MiniMax Hailuo 2.3 image-to-video generation — quick, cost-efficient cinematic clips from a single image.

    12 credits (discounted)#31 overall
    Hailuo

    Hailuo 02

    MiniMax's Hailuo 02 video generator — expressive, physical motion and strong prompt-to-action fidelity. Turns a plain-language action description into convincingly dynamic footage.

    32 credits#32 overall
    Luma

    Luma Ray 2 720p

    Luma Ray 2 video generation in 720p quality

    40 credits per second of output+1 variants
    Wan

    Wan Video 2.5 Image to Video

    High-quality image-to-video with Wan 2.5

    20 credits+2 variants
    ByteDance

    Seedance 1.5 Pro Text to Video

    Professional text-to-video with Seedance V1.5 Pro

    4 credits+4 variants
    ByteDance

    Seedance Image to Video

    Transform images into dynamic dancing videos

    3 credits+1 variants
    Pixverse

    Pixverse Text to Video

    Generate videos from text with Pixverse V5.5

    12 credits+1 variants
    Wan

    Wan Video 2.5 Text to Video

    High-quality text-to-video with Wan 2.5

    20 credits+3 variants
    Wan

    Wan 2.6 Image to Video

    Alibaba WAN 2.6 is a state-of-the-art image-to-video generation model with audio support, prompt expansion, and multiple resolution options

    20 credits (discounted)#42 overall
    Kling

    Kling V2.1

    Kling version 2.1 for advanced video generation

    56 credits (discounted)#43 overall
    LTX

    LTX 2.3 Text to Video Fast

    Fast text-to-video generation. Generates a video directly from a text prompt at 1080p/1440p/2160p, 6-20s duration, with optional native audio generation.

    20 credits#44 overall
    LTX

    LTX 2 Text to Video Pro

    Professional text-to-video with LTXV2

    29 credits#45 overall
    LTX

    LTX 2 Text to Video Fast

    Fast text-to-video generation with LTXV2

    20 credits#46 overall
    LTX

    LTX 2 Pro

    Professional image-to-video with LTXV2

    29 credits#47 overall
    LTX

    LTX 2

    Latest video generation with enhanced control

    20 credits#48 overall
    Midjourney

    Midjourney V7 Image to Video

    Midjourney V7 Image to Video generates high-quality videos from input images with cinematic motion and style consistency

    28 credits#54 overall
    LTX

    LTX 2.3 Image to Video Pro

    Pro-quality image-to-video generation. Animates a still image with a prompt. Supports 1080p/1440p/2160p, 6-20s duration (durations over 10s require 25fps + 1080p), optional audio generation and end-image transitions.

    29 credits#55 overall
    Runway

    Runway Text to Video

    Runway 4 for text to video generations

    72 credits+1 variants
    ByteDance

    ByteDance Upscaler Video

    AI-powered video upscaling from ByteDance

    3 credits
    NVIDIA

    Cosmos 3 Super Image to Video

    NVIDIA's Cosmos 3 Super world model animates a still image with physically grounded motion, guided by a text prompt that a built-in reasoner expands first. 480p or 720p, 2-7 seconds at 24 fps. Ranked in the top ten on the Artificial Analysis image-to-video board.

    8 credits
    ByteDance

    DreamActor V2

    AI-powered face reenactment model that transfers facial expressions and head movements from a driving video to a reference image

    20 credits
    Black Forest Labs

    Flux 3 Extend Video

    FLUX.3 continues an existing clip beyond its final frame, generating footage consistent with the original motion and scene (source clip must be MP4, under 50 MB and under 15 seconds)

    82 credits+1 variants
    Black Forest Labs

    Flux 3 First Last Frame to Video

    FLUX.3 interpolates a smooth, coherent transition between a defined start frame and end frame, with native audio — up to 1080p and 5–20 seconds

    34 credits
    Google

    Gemini Omni Flash Edit

    Google Gemini Omni Flash - conversational video editing (video-to-video)

    50 credits+3 variants
    Google

    Gemini Omni Video

    Google Gemini Omni multimodal video generation. Accepts a prompt plus optional reference images, source video clips, character IDs, and audio IDs. Supports 720p / 1080p / 4k output at 4-10 seconds (16:9 or 9:16). Quota: images + videos*2 + character_ids <= 7.

    29 credits
    Grok

    Grok Imagine Extend

    Grok Imagine video extension - continue an existing Grok Imagine video at a chosen timestamp, 6 or 10 seconds

    13 credits+1 variants
    Hailuo

    Hailuo 2.3 Standard

    Standard video generation with good quality

    23 credits
    Alibaba

    Happy Horse 1.0 Reference to Video

    Happy Horse 1.0 generates videos from up to 9 reference images using character1–character9 placeholders in the prompt, with native synchronized audio and multilingual lip-sync

    34 credits
    Alibaba

    Happy Horse 1.0 Video Edit

    Happy Horse 1.0 edits a source video using a prompt and up to 5 reference images, with auto/origin audio handling — output capped at 15 seconds

    56 credits
    Kling

    Kling 3 Turbo Text to Video

    Kling 3 Turbo fast text-to-video generation.

    22 credits (discounted)+1 variants
    Kling

    Kling O1 Image to Video

    Kling O1 image-to-video generation.

    45 credits+7 variants
    Kling

    Kling O3 Pro Image to Video

    Kling O3 Pro image-to-video with advanced reasoning-enhanced generation and camera controls

    56 credits
    Kling

    Kling O3 Standard Image to Video

    Kling O3 Standard image-to-video with reasoning-enhanced generation

    45 credits
    Kling

    Kling O3 Standard Reference to Video

    Kling O3 Standard reference-to-video for character-consistent generation

    45 credits
    Kling

    Kling Video V3 4K Text to Video

    Kling V3 4K text-to-video - cinema-grade resolution with advanced camera controls, audio generation, and multi-prompt support

    168 credits+1 variants
    Kling

    Kling Video V3 Pro Motion Control

    Kling V3 Pro motion control - transfer motion from a driving video to a reference image at pro-tier quality

    68 credits
    Kling

    Kling Video V3 Standard Motion Control

    Kling V3 Standard motion control - transfer motion from a driving video to a reference image

    51 credits
    LTX

    LTX 2 Retake Video

    Regenerate and improve video segments with LTX 2

    24 credits
    LTX

    LTX 2.3 Retake Video

    Retake a segment of an existing video. Takes a source video_url + prompt describing desired changes (e.g. "change flower to red rose"), plus optional start_time / duration / retake_mode (e.g. replace_audio_and_video).

    16 credits+2 variants
    LTX

    LTX 2.5 Text to Video Pro

    Quality-optimised text-to-video with synchronized native audio in a single pass, for final high-fidelity output. 720p or 1080p, up to 10 seconds, with eight camera-motion presets.

    58 credits+3 variants
    Decart

    Lucy Restyle

    AI-powered video restyling and transformation

    4 credits
    Pixverse

    Pixverse Effects

    Add special effects to videos with Pixverse

    12 credits
    Pixverse

    Pixverse Transition

    Create smooth transitions between video clips

    12 credits
    Sam

    Sam 3 Video Segment

    Intelligent video segmentation for object detection and precise video editing

    3 credits
    ByteDance

    Seedance 1 Pro Fast

    ByteDance Seedance 1 Pro Fast - accelerated image-to-video generation with high visual quality

    50 credits per second of output
    ByteDance

    Seedance 2.0 Fast Reference to Video

    Seedance 2.0 Fast Reference to Video generates cinematic videos from reference images, videos, and audio inputs with native audio-visual synchronization and director-level control — optimized for faster generation at lower cost

    55 credits (discounted)
    OpenAI

    Sora 2 Pro Storyboard

    Sora 2 with advanced storyboard scene generation

    240 credits
    OpenAI

    Sora 2 Text to Video

    Text-to-video generation with Sora 2

    32 credits+3 variants
    Topaz Labs

    Topaz Upscale Video

    Topaz Labs video upscaling — detail-preserving enhancement applied per frame, lifting soft or low-resolution footage toward crisp HD and beyond without the smeared look of naive scalers.

    4 credits (discounted)
    VEED

    VEED Video Background Removal

    VEED's video background removal — isolates the subject from moving footage frame by frame, no green screen required. Outputs a clean cutout ready for compositing onto any backdrop.

    9 credits+1 variants
    VEED

    VEED Video Background Removal Green Screen

    VEED background removal with a green-screen output — replaces the background with solid chroma green so you can key the footage in any editor downstream and keep full control of the composite.

    10 credits
    Google

    VEO 3.1 Extend Video

    Extend video duration with VEO 3.1

    112 credits+1 variants
    Google

    VEO 3.1 Reference to Video

    Google VEO 3.1 reference-to-video — generate an 8s cinematic clip whose subjects match up to 3 reference images, with native audio.

    128 credits
    Google

    VEO First Last Frame

    Generate videos between first and last frames

    64 credits+1 variants
    Other

    Video Enhancer Pro

    A heavyweight enhancement pass for finished footage — upscaling plus cleanup of compression artifacts, noise and softness in one run. The final polish before delivery.

    83 credits
    Vidu

    Vidu Q2 Pro Image to Video

    Vidu Q2 Pro animates a still image with precise duration control (any length from 2 to 8 seconds) and selectable motion intensity. 720p or 1080p; add an end frame for a transition. The higher-fidelity Q2 tier.

    16 credits+1 variants
    Wan

    Wan 2.1 Image to Video

    Wan 2.1 is an open-source AI video generation model that utilizes a diffusion transformer architecture and a novel 3D spatio-temporal VAE (Wan-VAE) for image-to-video generation.

    8 credits (discounted)+1 variants
    Wan

    Wan 2.2 14B Move

    Advanced video motion control with 14B parameters

    16 credits+1 variants
    Wan

    Wan 2.6 Image to Video Flash

    Ultra-fast image-to-video conversion with Wan 2.6 Flash technology

    20 credits
    Wan

    Wan 2.7 Reference to Video

    Wan 2.7 reference-to-video — generate videos from up to 5 reference images and/or videos, with optional first frame and voice timbre control

    20 credits
    Wan

    Wan 2.7 Video Edit

    Wan 2.7 video editing — modify a source video using prompts and an optional reference image for character, clothing, or style guidance

    20 credits
    Wan

    Wan 3.0 Prime Text to Video

    The accelerated tier of Wan 3.0: the same text-to-video model with native audio, served faster at a higher per-second rate. 480p, 720p or 1080p, 2-15 seconds, five aspect ratios or adaptive.

    11 credits+2 variants

    Image models

    66 pages

    Text-to-image, editing and vector output.

    ByteDanceFeatured

    Seedream 5.0 Pro

    ByteDance's flagship text-to-image model — deep-thinking prompt understanding, native text rendering in 14 languages, and precise control over dense layouts and structured designs

    6 credits per generation+3 variants
    MidjourneyFeatured

    Midjourney V7

    Midjourney V7 text-to-image generation with photorealistic and artistic outputs

    8 credits per generation#95 overall
    OpenAIFeatured

    GPT Image 2.5 Flare Edit

    Precise editing with OpenAI's default GPT Image 2.5 model: changes only what is asked, keeping subject, composition and background intact, with reference subjects staying recognisable across styles and successive edits. Up to 16 input images and an optional mask.

    5 credits+1 variants
    QwenFeatured

    Qwen Image 3 Text to Image

    Qwen's third-generation image model. Strong prompt adherence and complex text rendering, in both Chinese and English, with intelligent prompt rewriting on by default.

    3 credits per megapixel+3 variants
    OpenAI

    GPT Image 2 Text to Image

    OpenAI GPT Image 2 text-to-image. High-fidelity image generation with strong prompt adherence; supports up to 4K rendering.

    9 credits (discounted)+1 variants
    Google

    Nano Banana 2

    Ultra-fast lightweight image generation with Nano Banana 2 for rapid creative iteration

    4 credits per generation (discounted)+2 variants
    OpenAI

    GPT Image 1.5

    OpenAI's GPT Image 1.5 — conversational prompt understanding across text-to-image, image-to-image and edits. Strong at following long, specific instructions and rendering legible in-image text.

    2 credits per generation#6 overall
    Microsoft

    Mai Image 2.5 Edit

    Microsoft MAI-Image-2.5 image editing (pixel-level, cleanup, backgrounds, text)

    4 credits per generation+1 variants
    Grok

    Grok Imagine Image Quality

    xAI Grok Imagine Image (Quality tier) text-to-image. Quality-optimized generation up to 2K resolution with a wide range of aspect ratios.

    4 credits per generation#13 overall
    Luma

    Luma UNI 1 Max

    Luma Uni-1 (Max) — a reasoning-first image model that plans intent before generating, tuned for art-directed, cinematic results with optional reference images.

    1 credit per generation#16 overall
    Grok

    Grok Imagine Image

    xAI's Grok Imagine for still images — fast, loosely-styled generations with a distinctive creative bent. A strong ideation model when you want unexpected takes rather than strict prompt adherence.

    2 credits per generation#17 overall
    HiDream

    HiDream O1 Image

    HiDream O1 pixel-native text-to-image that reasons before it draws. Unified model producing high-resolution images up to 2K, with optional subject references for personalization.

    1 credit per megapixel#17 overall
    Flux

    Flux 2 Max

    Maximum quality image generation with the Flux 2 Max model for ultra-detailed photorealistic outputs

    6 credits per generation+3 variants
    ByteDance

    Seedream 4.0

    ByteDance Seedream 4.0 text-to-image. High-resolution generation (up to 4K) with multi-image output and built-in prompt enhancement.

    2 credits per generation (discounted)+1 variants
    ByteDance

    Seedream 4.5

    ByteDance's Seedream 4.5 all-round image model — generation, editing and image-to-image in one endpoint, with the strong text rendering and prompt comprehension the Seedream line is known for.

    4 credits per generation+2 variants
    Flux

    Flux 2 Flex

    Flexible image generation and editing model from the Flux 2 family with versatile creative controls

    4 credits per megapixel#22 overall
    Wan

    Wan 2.7 Pro Edit

    WAN 2.7 Pro delivers professional-grade image editing and transformation with text-guided instructions, supporting up to 4 reference images for precise modifications

    6 credits per image+3 variants
    Hunyuan

    HunyuanImage 3.0 Instruct Edit

    HunyuanImage 3.0 Instruct Edit is an advanced instruction-following image editing model by Tencent. It allows users to modify and transform images based on natural language instructions while preserving the original content and structure of the image.

    8 credits per megapixel+1 variants
    Ideogram

    Ideogram V4

    Ideogram V4 text-to-image. Creates high-quality images, posters, and logos with crisp visuals, accurate text rendering, fine detail, and full creative control.

    5 credits per generation#28 overall
    Kling

    Kling Image 3.0

    Kling V3 text-to-image generation with high-quality photorealistic and artistic outputs

    3 credits per generation#29 overall
    Flux

    Flux 2 Klein 9B Base

    Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

    1 credit per megapixel#30 overall
    Wan

    Wan 2.6 Text to Image

    Text-to-image generation with WAN V2.6

    2 credits per generation (discounted)+1 variants
    Krea

    Krea 2 Medium

    Krea 2 Medium text-to-image. Fast, photoreal generation with tunable creativity levels and optional style references.

    3 credits per generation+1 variants
    Qwen

    Qwen Image Edit 2511

    Qwen Image Edit 2511 is an image editing model from Alibaba that enables precise modifications to existing images based on textual instructions. It supports a wide range of editing tasks including object removal, style transfer, and content transformation.

    2 credits per megapixel (discounted)+1 variants
    PrunaAI

    Pruna Image Edit

    PrunaAI's efficiency-optimised image editor — instruction-based edits at the single-credit tier, compressed for speed. The economical pick for bulk retouch and iteration-heavy sessions.

    1 credit per generation (discounted)+1 variants
    Recraft

    Recraft 4.1 Text to Image

    Recraft V4.1 standard text-to-image. High-quality image generation from prompts with strong style control and consistency.

    3 credits per generation+5 variants
    Qwen

    Qwen Image Edit

    Alibaba Qwen's image editor — precise instruction-based edits with particular strength at changing or correcting text inside an image, a task most editors smear into noise.

    2 credits per megapixel (discounted)#42 overall
    NVIDIA

    Cosmos 3 Super

    NVIDIA Cosmos 3 Super text-to-image. High-fidelity generation with optional agentic refinement and prompt expansion for strong prompt adherence.

    4 credits per generation#43 overall
    HiDream

    HiDream O1 Image Edit

    HiDream O1 image editing. Instruction-driven edits and subject-driven changes on a supplied image at up to 2K resolution.

    1 credit per megapixel#45 overall
    Imagine

    ImagineArt 2.0

    ImagineArt 2.0 text-to-image with true-to-life realism, precise text rendering, and cinematic lighting. Supports 1K/2K output and low/high reasoning.

    4 credits per generation#46 overall
    Kling

    Kling Image O1

    Kling Image O1 is an advanced image generation model by Kuaishou that produces high-quality, detailed images from text prompts. It leverages state-of-the-art diffusion techniques to deliver visually compelling results for creative and professional use cases.

    3 credits per image#49 overall
    Flux

    Flux 2 Klein 4B Base Edit

    Flux 2 Klein 4B Base Edit is a compact yet powerful image editing model built on the Flux 2 architecture. It enables high-quality instruction-based image edits at an efficient parameter scale, making it suitable for fast and responsive image modification workflows.

    1 credit per megapixel#52 overall
    Recraft

    Recraft V4

    Latest Recraft V4 text-to-image model with exceptional design quality and typography support

    4 credits per generation#52 overall
    Flux

    Flux 2 Flash

    Ultra-fast image generation with Flux 2

    1 credit per megapixel+1 variants
    OpenAI

    GPT Image 1 Edit

    Instruction-based image editing with OpenAI's GPT Image — describe the change in plain English and it repaints only what you asked, preserving the rest of the frame. Suited to precise product and layout touch-ups.

    2 credits per generation#55 overall
    Imagine

    ImagineArt 1.5

    Imagine Art 1.5 — stylised, mood-led image generation that leans expressive rather than literal. Suits emotive social visuals, album-art energy and illustrative looks.

    4 credits per generation#61 overall
    ByteDance

    Seedream Edit

    ByteDance's Seedream editor — instruction-driven changes to an existing image: replace objects, adjust backgrounds, fix details. The budget-tier edit option in the Seedream family.

    2 credits per image (discounted)#62 overall
    Flux

    Flux Kontext

    Context-aware generation with coherent compositions

    2 credits per generation (discounted)+1 variants
    Z-Image

    Z Image Turbo

    Z-Image Turbo — an efficiency-tuned text-to-image model built for speed. Single-credit generations that hold up remarkably well for drafts, mood boards and rapid A/B exploration.

    1 credit per generation (discounted)#74 overall
    Google

    Imagen 4

    Premium model for photorealistic results (text-only)

    4 credits per generation+1 variants
    Baidu

    ERNIE Image Turbo

    Baidu ERNIE Image (Turbo) — an 8B diffusion-transformer distilled for fast generation, strong at posters, comics, and text-in-image layouts.

    1 credit per megapixel#78 overall
    Leonardo

    Leonardo Lucid Origin

    Leonardo's Lucid Origin — a versatile image model spanning photoreal to stylised output with strong adherence to detailed art direction. A dependable general-purpose pick for brand visuals.

    4 credits per generation#79 overall
    Reve

    Reve Text to Image

    Reve's image model — tight prompt adherence and clean, legible text rendering. Delivers composed, poster-ready frames from dense multi-element prompts.

    4 credits per generation#83 overall
    ByteDance

    ByteDance Seedream 4

    Seedream 4 — the previous-generation ByteDance image model. Still a capable, lower-cost text-to-image pick when you don't need 4.5's stronger prompt comprehension.

    3 credits per generation#87 overall
    Qwen

    Qwen Image

    High-quality realistic images with excellent detail

    1 credit per megapixel (discounted)+1 variants
    Flux

    Flux 1.1 Pro

    Professional grade Flux model with enhanced features

    4 credits per megapixel+2 variants
    HiDream

    HiDream O1 Image Dev

    HiDream O1 Image (Dev) — the faster, distilled tier of the pixel-native text-to-image model. Lower cost with strong prompt adherence up to 2K.

    1 credit per megapixel+1 variants
    Recraft

    Recraft V3 Image

    Recraft V3 — design-first image generation with clean vector-ready output, controlled brand styles and dependable typography. The pick for logos, icons and layout-quality graphics.

    4 credits per generation#107 overall
    Luma

    Luma Photon Flash

    Luma's Photon Flash — the speed tier of the Photon image family. Near-instant photorealistic stills at the lowest credit tier in the catalog, ideal for high-volume ideation and thumbnails.

    1 credit per megapixel#108 overall
    Qwen

    Qwen Z Image

    Advanced image generation model with enhanced creative capabilities

    2 credits per megapixel#112 overall
    Flux

    Flux Schnell

    Fast image generation model with high-quality outputs and efficient processing

    1 credit per megapixel (discounted)#134 overall
    Runway

    Runway Gen4 Image

    Runway Gen4 image generation model with high-quality creative outputs and style control

    4 credits+1 variants
    Clarity

    Clarity Crystal Upscaler

    Clarity's Crystal upscaler — sharpens and enlarges images while regenerating fine texture rather than stretching pixels. Rescues soft AI outputs and small source files for print or 4K use.

    2 credits per megapixel
    PrunaAI

    FireRed Image Edit

    Advanced image editing model for prompt-guided modifications, inpainting, and creative edits

    3 credits per megapixel
    OpenAI

    GPT Image 2.5 Sunburst Edit

    Editing with the tightest control from OpenAI's precision GPT Image 2.5 model: edits scoped exactly to the instruction, with subject and composition preserved across many rounds of revision. Up to 16 input images and an optional mask.

    5 credits+1 variants
    Grok

    Grok Imagine Image 2.0

    xAI Grok Imagine 2.0 text-to-image. Thirteen aspect ratios from 2:1 through 1:2, 1K or 2K output, and up to four images per request.

    5 credits+1 variants
    HiggsField

    Higgsfield

    HiggsField's premium still generator — high-drama, cinematic frames with pronounced lighting and motion energy. The high-impact option for hero shots and campaign key visuals.

    19 credits per generation
    Meituan

    LongCat Edit

    Longcat Edit is an image editing model that supports long-context image understanding and precise editing based on detailed textual instructions. It excels at complex multi-step edits and nuanced image transformations.

    12 credits per megapixel
    Midjourney

    Midjourney Niji 6

    Midjourney Niji 6 anime and illustration-focused image generation optimized for anime, manga, and cartoon styles

    8 credits per generation
    Google

    Nano Banana

    Fast and efficient image generation with vibrant colors

    4 credits per generation+3 variants
    Qwen

    Qwen Image 2 Edit

    Qwen Image 2 Edit is an image editing model that allows users to modify and enhance images based on specific instructions. It is designed to provide high-quality edits while maintaining the integrity of the original image, making it suitable for photo editing, retouching, and creative transformations.

    3 credits per megapixel+3 variants
    Recraft

    Recraft 4 Text to Image

    Recraft 4 is a powerful text-to-image model that generates high-quality images from textual descriptions. It is designed to create detailed and visually appealing images based on user input, making it ideal for various applications such as content creation, design, and more.

    4 credits per generation+3 variants
    Sourceful

    Riverflow 2.0

    Sourceful's Riverflow 2.0 Pro — production-grade image generation aimed at design and product work, where layout discipline and clean detailing matter more than raw style range.

    11 credits per generation
    Topaz Labs

    Topaz Image Colorization

    AI-powered colorization of black-and-white images with natural, realistic color restoration

    7 credits per generation
    Topaz Labs

    Topaz Upscale Image

    Topaz Labs' industry-standard image upscaler — detail-faithful enlargement to production resolution with strong artifact and noise cleanup. The trusted route from draft render to deliverable.

    7 credits per generation
    Wan

    Wan Preview 2.5 Text to Image

    Wan 2.5 in still-image mode — the video family's frame generator used standalone. Produces cinematic, video-native compositions that match Wan footage, ideal for first-frame and thumbnail work.

    4 credits per generation+1 variants

    Lipsync models

    16 pages

    Talking-head and avatar models that sync mouth to audio.

    MirageFeatured

    Avatar X Text to Video

    Type a script and a stock avatar performs it — voice included. Mirage's most advanced avatar model, with strong identity preservation and expressive delivery. Scripts run 50–1,500 characters.

    120 credits+1 variants
    HeyGenFeatured

    HeyGen Avatar V3

    Generate digital twin videos with HeyGen Avatar V3. Choose from 700+ premade avatars with customizable voices, expressions, and styles.

    14 credits
    HeyGenFeatured

    HeyGen Avatar V5

    HeyGen Avatar V digital twins — natural talking-avatar videos from a script or audio with premium lip-sync

    40 credits
    HeyGenFeatured

    HeyGen Image to Video

    Turn your photo into a talking avatar video. Upload a selfie and choose a voice — HeyGen Avatar 4 generates realistic lip-sync video.

    40 credits
    SyncFeatured

    Sync React 1

    Emotional lipsync with facial expressions and head movement. Supports emotions (happy, angry, sad, neutral, disgusted, surprised), model modes (lips, face, head), and temperature control.

    67 credits
    VEEDFeatured

    VEED Avatars

    Professional talking avatar generation with natural expressions and lip synchronization

    3 credits+1 variants
    Kling

    Kling Avatar Pro

    Professional quality talking avatar with advanced lip synchronization

    46 credits+1 variants
    Kling

    Kling Lipsync

    Kling's dedicated lipsync model — syncs a speaker's mouth in an existing video to new audio. Built for avatar animation and translated or re-voiced talking-head clips.

    6 credits per second of output
    LTX

    LTX 2 Audio to Video

    Generate videos from audio input with synchronized visuals and motion

    16 credits per generation
    LTX

    LTX 2.3 Audio to Video

    High-quality, fast AI video model for audio-to-video with lipsync, stylized transform capabilities

    40 credits
    LTX

    LTX 2.5 Audio to Video Pro

    Quality-optimised video timed to a supplied audio clip, for final visuals synchronized to music, dialogue or a soundtrack. Takes up to 10s of audio plus a prompt, an optional first-frame image, or both. Output length follows the audio.

    68 credits+1 variants
    Sync

    Sync Lipsync 2.0

    Professional video lipsync with sync mode control. Supports cut_off, loop, bounce, silence, and remap sync modes.

    20 credits+1 variants
    VEED

    VEED Fabric 1.0

    Professional fabric and texture video generation

    32 credits+1 variants
    VEED

    VEED Fabric 1.0 Text

    Turn a photo into a talking video from just a script — the voice is auto-generated to match the character

    32 credits
    VEED

    VEED Lipsync

    Fast and affordable video lipsync generation. Simple video + audio input.

    3 credits
    Wan

    Wan 2.2 Speech to Video

    Generates high-quality videos from static images and audio with realistic facial expressions, body movements, and professional camera work

    40 credits+1 variants

    Audio & voice models

    19 pages

    Speech, voice cloning and music generation.

    CartesiaFeatured

    Cartesia Sonic 3.5

    Cartesia's newest and top-ranked TTS model (Sonic 3.5). Multilingual, highly expressive, with emotion control, speed tuning, and voice cloning.

    4 credits per 1,000 characters+1 variants
    CartesiaFeatured

    Cartesia Sonic 3.6

    Cartesia's newest and top-ranked TTS model (Sonic 3.6). 44 languages including Hindi-English code-switching, context-driven emotion without tags, speed tuning, and voice cloning. Existing cloned voices work unchanged.

    4 credits per 1,000 characters
    InworldFeatured

    Inworld TTS 2

    Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language style steering for fine-grained delivery control.

    2 credits per 1,000 characters
    InworldFeatured

    Inworld TTS 2 Flash

    Inworld Realtime TTS-2 Flash: the low-latency tier of TTS-2, around five times faster to first audio and 40% cheaper, with the same 200+ languages and voice cloning. Best for high-volume and conversational use.

    2 credits per 1,000 characters
    ByteDanceFeatured

    Seed Audio 1.0

    ByteDance Seed Audio 1.0 — high-quality, natural-sounding text-to-speech with preset voices, optional reference audio (@Audio1–@Audio3) or a reference image, and speed/pitch/volume controls.

    17 credits per 1,000 characters
    Google

    Gemini 3.1 Flash TTS

    Google Gemini 3.1 Flash TTS — expressive text-to-speech with 30 voices, natural-language style control, inline audio tags ([sigh], [laughing], [whispering]), multilingual synthesis, and multi-speaker dialogue.

    12 credits per 1,000 characters#9 overall
    KIE

    ElevenLabs Multilingual

    Multilingual text-to-speech powered by ElevenLabs Multilingual v2.

    8 credits per 1,000 characters+1 variants
    MiniMax

    MiniMax Speech

    MiniMax's text-to-speech engine — clear multilingual delivery backed by one of the largest named-voice rosters in the catalog, so you can cast a specific voice rather than settle for a generic one.

    2 credits per 1,000 characters (discounted)#15 overall
    Chatterbox

    Chatterbox TTS

    Chatterbox text-to-speech — natural, expressive narration with inline vocal tags like <laugh> and <sigh> woven directly into the script. A solid default for UGC-style voiceovers.

    2 credits per 1,000 characters+1 variants
    Qwen

    Qwen 3 TTS 0.6B

    Compact text-to-speech model with natural voice synthesis and efficient processing

    2 credits per 1,000 characters#82 overall
    Qwen

    Qwen 3 TTS 1.7B

    High-quality text-to-speech model with enhanced naturalness and emotional expression

    2 credits per 1,000 characters#85 overall
    Cartesia

    Cartesia Voice Clone

    Clone any voice using Cartesia AI. Upload a short audio sample to instantly create a personalized voice for text-to-speech generation.

    8 credits
    Eleven Labs

    ElevenLabs Voice Change

    Voice conversion model that transforms voice characteristics while preserving speech content

    20 credits
    Grok

    Grok TTS

    xAI Grok TTS - high-quality text-to-speech with 5 expressive voices, 21 languages, and speech tags support. Up to 15,000 characters per request.

    4 credits per 1,000 characters
    Inworld

    Inworld Voice Clone

    Clone any voice using Inworld AI voice cloning. Upload audio samples to create a personalized voice for TTS generation.

    2 credits per 1,000 characters
    MiniMax

    MiniMax Music 3

    High-performance music generation for complete songs up to five minutes. Takes a style description plus lyrics, with structure tags such as [intro], [verse], [chorus] and [outro] on their own lines. Returns 44.1 kHz 16-bit stereo.

    3 credits per 1,000 characters
    Qwen

    Qwen 3 TTS Voice Design

    Qwen 3 text-to-speech with custom voice design

    161 credits per 1,000 characters
    Qwen

    Qwen Audio 3 TTS Flash

    Alibaba's hosted Qwen Audio 3.0 TTS (Flash tier): fast, natural multilingual speech across 10 languages with 44 named voices. The hosted successor to the open-weight Qwen 3 TTS models.

    4 credits per 1,000 characters
    Suno

    Suno Sounds V5.5

    Suno Sounds V5.5 is the latest sound generation model with improved quality, looping, tempo, key controls, and lyrics subtitle capture

    2 credits per generation+1 variants

    Video upscaling

    1 page

    Resolution and detail enhancement for footage you already have.

    Every model. One studio.

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.