AI Models

    296 AI models, one subscription.

    Every video, image, lipsync, upscaling and audio model Versely ships — real credit pricing, real rankings, no marketing fluff. Tier and input-mode variants of the same model share a page, so the 148 pages below cover the whole catalogue without repeating themselves.

    Browse by spec

    The same catalogue sliced by what you are actually choosing on — clip length, output quality, how a model bills, and what it can do. Each page prices the whole roster for that setting.

    Video models

    77 pages

    Text-to-video, image-to-video and video-to-video generation.

    AlibabaFeatured

    Happy Horse 1.0 Text to Video

    Happy Horse 1.0 generates expressive videos from text prompts with native synchronized audio, Foley sound effects, and multilingual lip-sync — supporting up to 1080p output and durations from 3 to 15 seconds

    28 credits+3 variants
    ByteDanceFeatured

    Seedance 2.0

    Seedance 2.0 generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability

    69 credits+7 variants
    WanFeatured

    Wan 2.7 Text to Video

    Wan 2.7 generates high-fidelity videos from text prompts with strong motion consistency, optional custom audio input, and intelligent prompt rewriting

    15 credits#6 overall
    AlibabaFeatured

    Happy Horse 1.1 Image to Video

    Happy Horse 1.1 animates a first-frame image into 1080p video with synchronized native audio and multilingual lip-sync (aspect ratio inferred from the image, 3–15s)

    18 credits#7 overall
    RunwayFeatured

    Runway Gen-4.5

    Runway Gen-4.5 video generation with state-of-the-art quality, cinematic motion, and precise camera controls

    50 credits per second of output#14 overall
    WanFeatured

    Wan 2.7 Image to Video

    Wan 2.7 animates images with three modes: first-frame to video, first-and-last-frame interpolation, or video continuation, with optional driving audio

    15 credits#17 overall
    PrunaAIFeatured

    PrunaAI P-Video

    Fast video generation with built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-to-video in a single endpoint. Generates 5s 720p video in ~10 seconds.

    2 credits per second of output#51 overall
    Black Forest LabsFeatured

    Flux 3 Text to Video

    FLUX.3 is Black Forest Labs' frontier video model — generates video with native audio directly from a text prompt, at up to 1080p and 5–20 seconds across eight aspect ratios

    9 credits per second of output
    MiniMaxFeatured

    Minimax H3 Text to Video

    MiniMax H3 is a frontier video model — generates 2K video from a text prompt alone, in durations from 5 to 15 seconds across six aspect ratios

    26 credits per second of output+2 variants
    Grok

    Grok Imagine Video

    Grok image-to-video generation

    7 credits#9 overall
    Kling

    Kling 2.5 Turbo

    Fast video generation with good quality balance

    21 credits#9 overall
    Vidu

    Vidu Q3 Image to Video

    Transform static images into dynamic videos with Vidu Q3 technology

    8 credits (discounted)#12 overall
    Vidu

    Vidu Q3 Video

    Generate high-quality videos from text prompts with Vidu Q3 engine

    16 credits#12 overall
    Pixverse

    Pixverse 5.6 Image to Video

    Latest Pixverse model for converting images to high-quality animated videos

    75 credits#14 overall
    Google

    VEO 3.1

    Enhanced VEO 3 with improved quality and features

    40 credits+3 variants
    Pixverse

    Pixverse 5.6 Text to Video

    Generate videos directly from text prompts with Pixverse 5.6 advanced capabilities

    75 credits#16 overall
    Kling

    Kling Video V3 Standard Image to Video

    Kling V3 Standard image-to-video generation with camera controls and element composition

    63 credits+3 variants
    Pixverse

    Pixverse V6 Text to Video

    PixVerse V6 text-to-video. Generates video from a text prompt with strong motion and prompt control, optional audio and multi-clip, up to 1080p.

    9 credits+1 variants
    Pixverse

    Pixverse V4

    Pixverse version 4 for creative video generation

    40 credits#20 overall
    Kling

    Kling Video V2.6 Pro Text to Video

    Professional text-to-video with Kling V2.6 Pro

    70 credits+3 variants
    Kling

    Kling O3 Pro Text to Video

    Kling O3 Pro text-to-video with reasoning-enhanced generation, camera controls, and audio

    70 credits+6 variants
    Wan

    Wan V2.6 Text to Video

    Text-to-video generation with WAN V2.6

    15 credits+2 variants
    Hailuo

    Hailuo 2.3 Pro

    Professional-grade video generation with enhanced features

    49 credits#29 overall
    Luma

    Luma Ray 2 720p

    Luma Ray 2 video generation in 720p quality

    50 credits per second of output+1 variants
    Hailuo

    Hailuo 2.3 Fast

    Fast MiniMax Hailuo 2.3 image-to-video generation — quick, cost-efficient cinematic clips from a single image.

    14 credits (discounted)#31 overall
    ByteDance

    Seedance 1.5 Pro Text to Video

    Professional text-to-video with Seedance V1.5 Pro

    6 credits+4 variants
    ByteDance

    Seedance Image to Video

    Transform images into dynamic dancing videos

    9 credits+1 variants
    Pixverse

    Pixverse Text to Video

    Generate videos from text with Pixverse V5.5

    40 credits+1 variants
    Wan

    Wan 2.6 Image to Video

    Alibaba WAN 2.6 is a state-of-the-art image-to-video generation model with audio support, prompt expansion, and multiple resolution options

    8 credits (discounted)#39 overall
    Kling

    Kling V2.1

    Kling version 2.1 for advanced video generation

    70 credits (discounted)#40 overall
    LTX

    LTX 2.3 Text to Video Fast

    Fast text-to-video generation. Generates a video directly from a text prompt at 1080p/1440p/2160p, 6-20s duration, with optional native audio generation.

    4 credits#41 overall
    LTX

    LTX 2 Text to Video Pro

    Professional text-to-video with LTXV2

    6 credits#42 overall
    LTX

    LTX 2 Pro

    Professional image-to-video with LTXV2

    6 credits#43 overall
    LTX

    LTX 2 Text to Video Fast

    Fast text-to-video generation with LTXV2

    4 credits#43 overall
    LTX

    LTX 2

    Latest video generation with enhanced control

    4 credits#45 overall
    Midjourney

    Midjourney V7 Image to Video

    Midjourney V7 Image to Video generates high-quality videos from input images with cinematic motion and style consistency

    110 credits#49 overall
    LTX

    LTX 2.3 Image to Video Pro

    Pro-quality image-to-video generation. Animates a still image with a prompt. Supports 1080p/1440p/2160p, 6-20s duration (durations over 10s require 25fps + 1080p), optional audio generation and end-image transitions.

    6 credits#51 overall
    Runway

    Runway Text to Video

    Runway 4 for text to video generations

    18 credits per second of output+1 variants
    Bria

    Bria Video Background Removal

    AI-powered video background removal and replacement

    1 credit
    ByteDance

    ByteDance Upscaler Video

    AI-powered video upscaling from ByteDance

    1 credit
    ByteDance

    DreamActor V2

    AI-powered face reenactment model that transfers facial expressions and head movements from a driving video to a reference image

    5 credits per second of output
    Black Forest Labs

    Flux 3 Extend Video

    FLUX.3 continues an existing clip beyond its final frame, generating footage consistent with the original motion and scene (source clip must be MP4, under 50 MB and under 15 seconds)

    21 credits per second of output+1 variants
    Black Forest Labs

    Flux 3 First Last Frame to Video

    FLUX.3 interpolates a smooth, coherent transition between a defined start frame and end frame, with native audio — up to 1080p and 5–20 seconds

    9 credits per second of output
    Google

    Gemini Omni Flash Edit

    Google Gemini Omni Flash - conversational video editing (video-to-video)

    13 credits per second of output+3 variants
    Google

    Gemini Omni Video

    Google Gemini Omni multimodal video generation. Accepts a prompt plus optional reference images, source video clips, character IDs, and audio IDs. Supports 720p / 1080p / 4k output at 4-10 seconds (16:9 or 9:16). Quota: images + videos*2 + character_ids <= 7.

    9 credits per second of output
    Grok

    Grok Imagine Extend

    Grok Imagine video extension - continue an existing Grok Imagine video at a chosen timestamp, 6 or 10 seconds

    8 credits+1 variants
    Hailuo

    Hailuo 2.3 Standard

    Standard video generation with good quality

    28 credits
    Alibaba

    Happy Horse 1.0 Reference to Video

    Happy Horse 1.0 generates videos from up to 9 reference images using character1–character9 placeholders in the prompt, with native synchronized audio and multilingual lip-sync

    28 credits
    Alibaba

    Happy Horse 1.0 Video Edit

    Happy Horse 1.0 edits a source video using a prompt and up to 5 reference images, with auto/origin audio handling — output capped at 15 seconds

    28 credits
    Kling

    Kling 3 Turbo Text to Video

    Kling 3 Turbo fast text-to-video generation.

    10 credits per second of output (discounted)+1 variants
    Kling

    Kling O1 Image to Video

    Kling O1 image-to-video generation.

    56 credits+7 variants
    Kling

    Kling O3 Pro Image to Video

    Kling O3 Pro image-to-video with advanced reasoning-enhanced generation and camera controls

    70 credits
    Kling

    Kling O3 Standard Image to Video

    Kling O3 Standard image-to-video with reasoning-enhanced generation

    56 credits
    Kling

    Kling O3 Standard Reference to Video

    Kling O3 Standard reference-to-video for character-consistent generation

    56 credits
    Kling

    Kling Video V3 4K Text to Video

    Kling V3 4K text-to-video - cinema-grade resolution with advanced camera controls, audio generation, and multi-prompt support

    210 credits+1 variants
    Kling

    Kling Video V3 Pro Motion Control

    Kling V3 Pro motion control - transfer motion from a driving video to a reference image at pro-tier quality

    17 credits per second of output
    Kling

    Kling Video V3 Standard Motion Control

    Kling V3 Standard motion control - transfer motion from a driving video to a reference image

    13 credits per second of output
    LTX

    LTX 2 Retake Video

    Regenerate and improve video segments with LTX 2

    10 credits per second of output
    LTX

    LTX 2.3 Retake Video

    Retake a segment of an existing video. Takes a source video_url + prompt describing desired changes (e.g. "change flower to red rose"), plus optional start_time / duration / retake_mode (e.g. replace_audio_and_video).

    10 credits per second of output+2 variants
    Decart

    Lucy Restyle

    AI-powered video restyling and transformation

    1 credit per second of output
    Pixverse

    Pixverse Effects

    Add special effects to videos with Pixverse

    40 credits
    Pixverse

    Pixverse Transition

    Create smooth transitions between video clips

    40 credits
    Sam

    Sam 3 Video Segment

    Intelligent video segmentation for object detection and precise video editing

    1 credit per second of output
    ByteDance

    Seedance 1 Pro Fast

    ByteDance Seedance 1 Pro Fast - accelerated image-to-video generation with high visual quality

    62 credits per second of output
    ByteDance

    Seedance 2.0 Fast Reference to Video

    Seedance 2.0 Fast Reference to Video generates cinematic videos from reference images, videos, and audio inputs with native audio-visual synchronization and director-level control — optimized for faster generation at lower cost

    18 credits per second of output (discounted)
    OpenAI

    Sora 2 Pro Storyboard

    Sora 2 with advanced storyboard scene generation

    30 credits per second of output
    OpenAI

    Sora 2 Text to Video

    Text-to-video generation with Sora 2

    10 credits per second of output+3 variants
    Google

    VEO 3.1 Extend Video

    Extend video duration with VEO 3.1

    40 credits+1 variants
    Google

    VEO 3.1 Reference to Video

    Google VEO 3.1 reference-to-video — generate an 8s cinematic clip whose subjects match up to 3 reference images, with native audio.

    64 credits
    Google

    VEO First Last Frame

    Generate videos between first and last frames

    40 credits+1 variants
    Wan

    Wan 2.1 Image to Video

    Wan 2.1 is an open-source AI video generation model that utilizes a diffusion transformer architecture and a novel 3D spatio-temporal VAE (Wan-VAE) for image-to-video generation.

    40 credits (discounted)+1 variants
    Wan

    Wan 2.2 14B Move

    Advanced video motion control with 14B parameters

    8 credits+1 variants
    Wan

    Wan 2.6 Image to Video Flash

    Ultra-fast image-to-video conversion with Wan 2.6 Flash technology

    8 credits
    Wan

    Wan 2.7 Reference to Video

    Wan 2.7 reference-to-video — generate videos from up to 5 reference images and/or videos, with optional first frame and voice timbre control

    18 credits
    Wan

    Wan 2.7 Video Edit

    Wan 2.7 video editing — modify a source video using prompts and an optional reference image for character, clothing, or style guidance

    18 credits
    Wan

    Wan Video 2.5 Image to Video

    High-quality image-to-video with Wan 2.5

    15 credits
    Wan

    Wan Video 2.5 Text to Video

    High-quality text-to-video with Wan 2.5

    15 credits+2 variants

    Image models

    44 pages

    Text-to-image, editing and vector output.

    ByteDanceFeatured

    Seedream 5.0 Pro

    ByteDance's flagship text-to-image model — deep-thinking prompt understanding, native text rendering in 14 languages, and precise control over dense layouts and structured designs

    7 credits per generation+3 variants
    MidjourneyFeatured

    Midjourney V7

    Midjourney V7 text-to-image generation with photorealistic and artistic outputs

    10 credits per generation#84 overall
    OpenAI

    GPT Image 2 Text to Image

    OpenAI GPT Image 2 text-to-image. High-fidelity image generation with strong prompt adherence; supports up to 4K rendering.

    2 credits (discounted)+1 variants
    Microsoft

    Mai Image 2.5 Edit

    Microsoft MAI-Image-2.5 image editing (pixel-level, cleanup, backgrounds, text)

    5 credits per generation+1 variants
    Google

    Nano Banana 2

    Ultra-fast lightweight image generation with Nano Banana 2 for rapid creative iteration

    4 credits per generation (discounted)+2 variants
    Grok

    Grok Imagine Image Quality

    xAI Grok Imagine Image (Quality tier) text-to-image. Quality-optimized generation up to 2K resolution with a wide range of aspect ratios.

    5 credits per generation#9 overall
    Hunyuan

    HunyuanImage 3.0 Instruct Edit

    HunyuanImage 3.0 Instruct Edit is an advanced instruction-following image editing model by Tencent. It allows users to modify and transform images based on natural language instructions while preserving the original content and structure of the image.

    9 credits per megapixel+1 variants
    HiDream

    HiDream O1 Image

    HiDream O1 pixel-native text-to-image that reasons before it draws. Unified model producing high-resolution images up to 2K, with optional subject references for personalization.

    1 credit per megapixel#11 overall
    Luma

    Luma UNI 1 Max

    Luma Uni-1 (Max) — a reasoning-first image model that plans intent before generating, tuned for art-directed, cinematic results with optional reference images.

    1 credit per generation#11 overall
    Flux

    Flux 2 Max

    Maximum quality image generation with the Flux 2 Max model for ultra-detailed photorealistic outputs

    7 credits per generation+3 variants
    Kling

    Kling Image 3.0

    Kling V3 text-to-image generation with high-quality photorealistic and artistic outputs

    3 credits per generation#13 overall
    Kling

    Kling Image O1

    Kling Image O1 is an advanced image generation model by Kuaishou that produces high-quality, detailed images from text prompts. It leverages state-of-the-art diffusion techniques to deliver visually compelling results for creative and professional use cases.

    3 credits per image#16 overall
    ByteDance

    Seedream 4.0

    ByteDance Seedream 4.0 text-to-image. High-resolution generation (up to 4K) with multi-image output and built-in prompt enhancement.

    2 credits per generation (discounted)+1 variants
    Wan

    Wan 2.7 Pro Edit

    WAN 2.7 Pro delivers professional-grade image editing and transformation with text-guided instructions, supporting up to 4 reference images for precise modifications

    8 credits per image+3 variants
    Flux

    Flux 2 Flex

    Flexible image generation and editing model from the Flux 2 family with versatile creative controls

    5 credits per megapixel#19 overall
    HiDream

    HiDream O1 Image Edit

    HiDream O1 image editing. Instruction-driven edits and subject-driven changes on a supplied image at up to 2K resolution.

    1 credit per megapixel#19 overall
    Ideogram

    Ideogram V4

    Ideogram V4 text-to-image. Creates high-quality images, posters, and logos with crisp visuals, accurate text rendering, fine detail, and full creative control.

    6 credits per generation#24 overall
    Krea

    Krea 2 Medium

    Krea 2 Medium text-to-image. Fast, photoreal generation with tunable creativity levels and optional style references.

    3 credits per generation+1 variants
    Flux

    Flux 2 Klein 9B Base

    Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

    2 credits per megapixel#29 overall
    Wan

    Wan 2.6 Text to Image

    Text-to-image generation with WAN V2.6

    2 credits per generation (discounted)+1 variants
    Qwen

    Qwen Image Edit 2511

    Qwen Image Edit 2511 is an image editing model from Alibaba that enables precise modifications to existing images based on textual instructions. It supports a wide range of editing tasks including object removal, style transfer, and content transformation.

    2 credits per megapixel (discounted)+1 variants
    HiDream

    HiDream O1 Image Dev

    HiDream O1 Image (Dev) — the faster, distilled tier of the pixel-native text-to-image model. Lower cost with strong prompt adherence up to 2K.

    1 credit per megapixel+1 variants
    Baidu

    ERNIE Image Turbo

    Baidu ERNIE Image (Turbo) — an 8B diffusion-transformer distilled for fast generation, strong at posters, comics, and text-in-image layouts.

    1 credit per megapixel#34 overall
    Imagine

    ImagineArt 2.0

    ImagineArt 2.0 text-to-image with true-to-life realism, precise text rendering, and cinematic lighting. Supports 1K/2K output and low/high reasoning.

    5 credits per generation#37 overall
    Recraft

    Recraft 4.1 Text to Image

    Recraft V4.1 standard text-to-image. High-quality image generation from prompts with strong style control and consistency.

    4 credits per generation+5 variants
    Flux

    Flux 2 Flash

    Ultra-fast image generation with Flux 2

    1 credit per megapixel+1 variants
    Recraft

    Recraft V4

    Latest Recraft V4 text-to-image model with exceptional design quality and typography support

    5 credits per generation#45 overall
    Flux

    Flux 2 Klein 4B Base Edit

    Flux 2 Klein 4B Base Edit is a compact yet powerful image editing model built on the Flux 2 architecture. It enables high-quality instruction-based image edits at an efficient parameter scale, making it suitable for fast and responsive image modification workflows.

    1 credit per megapixel#49 overall
    Flux

    Flux Kontext

    Context-aware generation with coherent compositions

    2 credits per generation (discounted)+1 variants
    Google

    Imagen 4

    Premium model for photorealistic results (text-only)

    4 credits per generation+2 variants
    Qwen

    Qwen Image

    High-quality realistic images with excellent detail

    1 credit per megapixel (discounted)+1 variants
    Flux

    Flux 1.1 Pro

    Professional grade Flux model with enhanced features

    4 credits per megapixel+2 variants
    Qwen

    Qwen Z Image

    Advanced image generation model with enhanced creative capabilities

    2 credits per megapixel#103 overall
    Flux

    Flux Schnell

    Fast image generation model with high-quality outputs and efficient processing

    1 credit per megapixel (discounted)#121 overall
    Runway

    Runway Gen4 Image

    Runway Gen4 image generation model with high-quality creative outputs and style control

    8 credits+1 variants
    NVIDIA

    Cosmos 3 Super

    NVIDIA Cosmos 3 Super text-to-image. High-fidelity generation with optional agentic refinement and prompt expansion for strong prompt adherence.

    4 credits per generation
    PrunaAI

    FireRed Image Edit

    Advanced image editing model for prompt-guided modifications, inpainting, and creative edits

    4 credits per megapixel
    Meituan

    LongCat Edit

    Longcat Edit is an image editing model that supports long-context image understanding and precise editing based on detailed textual instructions. It excels at complex multi-step edits and nuanced image transformations.

    15 credits per megapixel
    Midjourney

    Midjourney Niji 6

    Midjourney Niji 6 anime and illustration-focused image generation optimized for anime, manga, and cartoon styles

    10 credits per generation
    Google

    Nano Banana

    Fast and efficient image generation with vibrant colors

    4 credits per generation+3 variants
    Qwen

    Qwen Image 2 Edit

    Qwen Image 2 Edit is an image editing model that allows users to modify and enhance images based on specific instructions. It is designed to provide high-quality edits while maintaining the integrity of the original image, making it suitable for photo editing, retouching, and creative transformations.

    4 credits per megapixel+3 variants
    Recraft

    Recraft 4 Text to Image

    Recraft 4 is a powerful text-to-image model that generates high-quality images from textual descriptions. It is designed to create detailed and visually appealing images based on user input, making it ideal for various applications such as content creation, design, and more.

    4 credits per generation+3 variants
    SeedVR

    SeedVR Upscale

    VR-optimized image upscaling

    4 credits
    Topaz Labs

    Topaz Image Colorization

    AI-powered colorization of black-and-white images with natural, realistic color restoration

    8 credits per generation

    Lipsync models

    13 pages

    Talking-head and avatar models that sync mouth to audio.

    HeyGenFeatured

    HeyGen Avatar V3

    Generate digital twin videos with HeyGen Avatar V3. Choose from 700+ premade avatars with customizable voices, expressions, and styles.

    17 credits
    HeyGenFeatured

    HeyGen Avatar V5

    HeyGen Avatar V digital twins — natural talking-avatar videos from a script or audio with premium lip-sync

    50 credits
    HeyGenFeatured

    HeyGen Image to Video

    Turn your photo into a talking avatar video. Upload a selfie and choose a voice — HeyGen Avatar 4 generates realistic lip-sync video.

    50 credits
    SyncFeatured

    Sync React 1

    Emotional lipsync with facial expressions and head movement. Supports emotions (happy, angry, sad, neutral, disgusted, surprised), model modes (lips, face, head), and temperature control.

    84 credits
    VEEDFeatured

    VEED Avatars

    Professional talking avatar generation with natural expressions and lip synchronization

    30 credits+1 variants
    Kling

    Kling Avatar Pro

    Professional quality talking avatar with advanced lip synchronization

    58 credits+1 variants
    LTX

    LTX 2 Audio to Video

    Generate videos from audio input with synchronized visuals and motion

    20 credits per generation
    LTX

    LTX 2.3 Audio to Video

    High-quality, fast AI video model for audio-to-video with lipsync, stylized transform capabilities

    20 credits
    Sync

    Sync Lipsync 2.0

    Professional video lipsync with sync mode control. Supports cut_off, loop, bounce, silence, and remap sync modes.

    25 credits+1 variants
    VEED

    VEED Fabric 1.0

    Professional fabric and texture video generation

    40 credits+1 variants
    VEED

    VEED Fabric 1.0 Text

    Turn a photo into a talking video from just a script — the voice is auto-generated to match the character

    40 credits
    VEED

    VEED Lipsync

    Fast and affordable video lipsync generation. Simple video + audio input.

    4 credits
    Wan

    Wan 2.2 Speech to Video

    Generates high-quality videos from static images and audio with realistic facial expressions, body movements, and professional camera work

    50 credits+1 variants

    Audio & voice models

    13 pages

    Speech, voice cloning and music generation.

    CartesiaFeatured

    Cartesia Sonic 3.5

    Cartesia's newest and top-ranked TTS model (Sonic 3.5). Multilingual, highly expressive, with emotion control, speed tuning, and voice cloning.

    5 credits+1 variants
    InworldFeatured

    Inworld TTS 1.5 Max

    Inworld Realtime TTS 1.5 Max — the #1 ranked Inworld model, delivering the best balance of quality and speed. Expressive, multilingual (15 languages), with voice library support.

    4 credits#7 overall
    InworldFeatured

    Inworld TTS 2

    Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language style steering for fine-grained delivery control.

    5 credits#9 overall
    ByteDanceFeatured

    Seed Audio 1.0

    ByteDance Seed Audio 1.0 — high-quality, natural-sounding text-to-speech with preset voices, optional reference audio (@Audio1–@Audio3) or a reference image, and speed/pitch/volume controls.

    4 credits
    Google

    Gemini 3.1 Flash TTS

    Google Gemini 3.1 Flash TTS — expressive text-to-speech with 30 voices, natural-language style control, inline audio tags ([sigh], [laughing], [whispering]), multilingual synthesis, and multi-speaker dialogue.

    4 credits#3 overall
    KIE

    ElevenLabs Multilingual

    Multilingual text-to-speech powered by ElevenLabs Multilingual v2.

    6 credits+1 variants
    Qwen

    Qwen 3 TTS 0.6B

    Compact text-to-speech model with natural voice synthesis and efficient processing

    2 credits+1 variants
    Cartesia

    Cartesia Voice Clone

    Clone any voice using Cartesia AI. Upload a short audio sample to instantly create a personalized voice for text-to-speech generation.

    8 credits
    Eleven Labs

    ElevenLabs Voice Change

    Voice conversion model that transforms voice characteristics while preserving speech content

    5 credits per second of output
    Grok

    Grok TTS

    xAI Grok TTS - high-quality text-to-speech with 5 expressive voices, 21 languages, and speech tags support. Up to 15,000 characters per request.

    4 credits
    Inworld

    Inworld TTS

    High-quality text-to-speech powered by Inworld AI. Supports a curated library of expressive voices with multilingual capability.

    3 credits+1 variants
    Qwen

    Qwen 3 TTS Voice Design

    Qwen 3 text-to-speech with custom voice design

    3 credits
    Suno

    Suno Sounds V5.5

    Suno Sounds V5.5 is the latest sound generation model with improved quality, looping, tempo, key controls, and lyrics subtitle capture

    2 credits per generation+1 variants

    Video upscaling

    1 page

    Resolution and detail enhancement for footage you already have.

    Every model. One app.

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.