Browse by spec
The same catalogue sliced by what you are actually choosing on — clip length, output quality, how a model bills, and what it can do. Each page prices the whole roster for that setting.
Video models
77 pagesText-to-video, image-to-video and video-to-video generation.
Happy Horse 1.0 Text to Video
Happy Horse 1.0 generates expressive videos from text prompts with native synchronized audio, Foley sound effects, and multilingual lip-sync — supporting up to 1080p output and durations from 3 to 15 seconds
Seedance 2.0
Seedance 2.0 generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability
Wan 2.7 Text to Video
Wan 2.7 generates high-fidelity videos from text prompts with strong motion consistency, optional custom audio input, and intelligent prompt rewriting
Happy Horse 1.1 Image to Video
Happy Horse 1.1 animates a first-frame image into 1080p video with synchronized native audio and multilingual lip-sync (aspect ratio inferred from the image, 3–15s)
Runway Gen-4.5
Runway Gen-4.5 video generation with state-of-the-art quality, cinematic motion, and precise camera controls
Wan 2.7 Image to Video
Wan 2.7 animates images with three modes: first-frame to video, first-and-last-frame interpolation, or video continuation, with optional driving audio
PrunaAI P-Video
Fast video generation with built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-to-video in a single endpoint. Generates 5s 720p video in ~10 seconds.
Flux 3 Text to Video
FLUX.3 is Black Forest Labs' frontier video model — generates video with native audio directly from a text prompt, at up to 1080p and 5–20 seconds across eight aspect ratios
Minimax H3 Text to Video
MiniMax H3 is a frontier video model — generates 2K video from a text prompt alone, in durations from 5 to 15 seconds across six aspect ratios
Grok Imagine Video
Grok image-to-video generation
Kling 2.5 Turbo
Fast video generation with good quality balance
Vidu Q3 Image to Video
Transform static images into dynamic videos with Vidu Q3 technology
Vidu Q3 Video
Generate high-quality videos from text prompts with Vidu Q3 engine
Pixverse 5.6 Image to Video
Latest Pixverse model for converting images to high-quality animated videos
VEO 3.1
Enhanced VEO 3 with improved quality and features
Pixverse 5.6 Text to Video
Generate videos directly from text prompts with Pixverse 5.6 advanced capabilities
Kling Video V3 Standard Image to Video
Kling V3 Standard image-to-video generation with camera controls and element composition
Pixverse V6 Text to Video
PixVerse V6 text-to-video. Generates video from a text prompt with strong motion and prompt control, optional audio and multi-clip, up to 1080p.
Pixverse V4
Pixverse version 4 for creative video generation
Kling Video V2.6 Pro Text to Video
Professional text-to-video with Kling V2.6 Pro
Kling O3 Pro Text to Video
Kling O3 Pro text-to-video with reasoning-enhanced generation, camera controls, and audio
Wan V2.6 Text to Video
Text-to-video generation with WAN V2.6
Hailuo 2.3 Pro
Professional-grade video generation with enhanced features
Luma Ray 2 720p
Luma Ray 2 video generation in 720p quality
Hailuo 2.3 Fast
Fast MiniMax Hailuo 2.3 image-to-video generation — quick, cost-efficient cinematic clips from a single image.
Seedance 1.5 Pro Text to Video
Professional text-to-video with Seedance V1.5 Pro
Seedance Image to Video
Transform images into dynamic dancing videos
Pixverse Text to Video
Generate videos from text with Pixverse V5.5
Wan 2.6 Image to Video
Alibaba WAN 2.6 is a state-of-the-art image-to-video generation model with audio support, prompt expansion, and multiple resolution options
Kling V2.1
Kling version 2.1 for advanced video generation
LTX 2.3 Text to Video Fast
Fast text-to-video generation. Generates a video directly from a text prompt at 1080p/1440p/2160p, 6-20s duration, with optional native audio generation.
LTX 2 Text to Video Pro
Professional text-to-video with LTXV2
LTX 2 Pro
Professional image-to-video with LTXV2
LTX 2 Text to Video Fast
Fast text-to-video generation with LTXV2
LTX 2
Latest video generation with enhanced control
Midjourney V7 Image to Video
Midjourney V7 Image to Video generates high-quality videos from input images with cinematic motion and style consistency
LTX 2.3 Image to Video Pro
Pro-quality image-to-video generation. Animates a still image with a prompt. Supports 1080p/1440p/2160p, 6-20s duration (durations over 10s require 25fps + 1080p), optional audio generation and end-image transitions.
Runway Text to Video
Runway 4 for text to video generations
Bria Video Background Removal
AI-powered video background removal and replacement
ByteDance Upscaler Video
AI-powered video upscaling from ByteDance
DreamActor V2
AI-powered face reenactment model that transfers facial expressions and head movements from a driving video to a reference image
Flux 3 Extend Video
FLUX.3 continues an existing clip beyond its final frame, generating footage consistent with the original motion and scene (source clip must be MP4, under 50 MB and under 15 seconds)
Flux 3 First Last Frame to Video
FLUX.3 interpolates a smooth, coherent transition between a defined start frame and end frame, with native audio — up to 1080p and 5–20 seconds
Gemini Omni Flash Edit
Google Gemini Omni Flash - conversational video editing (video-to-video)
Gemini Omni Video
Google Gemini Omni multimodal video generation. Accepts a prompt plus optional reference images, source video clips, character IDs, and audio IDs. Supports 720p / 1080p / 4k output at 4-10 seconds (16:9 or 9:16). Quota: images + videos*2 + character_ids <= 7.
Grok Imagine Extend
Grok Imagine video extension - continue an existing Grok Imagine video at a chosen timestamp, 6 or 10 seconds
Hailuo 2.3 Standard
Standard video generation with good quality
Happy Horse 1.0 Reference to Video
Happy Horse 1.0 generates videos from up to 9 reference images using character1–character9 placeholders in the prompt, with native synchronized audio and multilingual lip-sync
Happy Horse 1.0 Video Edit
Happy Horse 1.0 edits a source video using a prompt and up to 5 reference images, with auto/origin audio handling — output capped at 15 seconds
Kling 3 Turbo Text to Video
Kling 3 Turbo fast text-to-video generation.
Kling O1 Image to Video
Kling O1 image-to-video generation.
Kling O3 Pro Image to Video
Kling O3 Pro image-to-video with advanced reasoning-enhanced generation and camera controls
Kling O3 Standard Image to Video
Kling O3 Standard image-to-video with reasoning-enhanced generation
Kling O3 Standard Reference to Video
Kling O3 Standard reference-to-video for character-consistent generation
Kling Video V3 4K Text to Video
Kling V3 4K text-to-video - cinema-grade resolution with advanced camera controls, audio generation, and multi-prompt support
Kling Video V3 Pro Motion Control
Kling V3 Pro motion control - transfer motion from a driving video to a reference image at pro-tier quality
Kling Video V3 Standard Motion Control
Kling V3 Standard motion control - transfer motion from a driving video to a reference image
LTX 2 Retake Video
Regenerate and improve video segments with LTX 2
LTX 2.3 Retake Video
Retake a segment of an existing video. Takes a source video_url + prompt describing desired changes (e.g. "change flower to red rose"), plus optional start_time / duration / retake_mode (e.g. replace_audio_and_video).
Lucy Restyle
AI-powered video restyling and transformation
Pixverse Effects
Add special effects to videos with Pixverse
Pixverse Transition
Create smooth transitions between video clips
Sam 3 Video Segment
Intelligent video segmentation for object detection and precise video editing
Seedance 1 Pro Fast
ByteDance Seedance 1 Pro Fast - accelerated image-to-video generation with high visual quality
Seedance 2.0 Fast Reference to Video
Seedance 2.0 Fast Reference to Video generates cinematic videos from reference images, videos, and audio inputs with native audio-visual synchronization and director-level control — optimized for faster generation at lower cost
Sora 2 Pro Storyboard
Sora 2 with advanced storyboard scene generation
Sora 2 Text to Video
Text-to-video generation with Sora 2
VEO 3.1 Extend Video
Extend video duration with VEO 3.1
VEO 3.1 Reference to Video
Google VEO 3.1 reference-to-video — generate an 8s cinematic clip whose subjects match up to 3 reference images, with native audio.
VEO First Last Frame
Generate videos between first and last frames
Wan 2.1 Image to Video
Wan 2.1 is an open-source AI video generation model that utilizes a diffusion transformer architecture and a novel 3D spatio-temporal VAE (Wan-VAE) for image-to-video generation.
Wan 2.2 14B Move
Advanced video motion control with 14B parameters
Wan 2.6 Image to Video Flash
Ultra-fast image-to-video conversion with Wan 2.6 Flash technology
Wan 2.7 Reference to Video
Wan 2.7 reference-to-video — generate videos from up to 5 reference images and/or videos, with optional first frame and voice timbre control
Wan 2.7 Video Edit
Wan 2.7 video editing — modify a source video using prompts and an optional reference image for character, clothing, or style guidance
Wan Video 2.5 Image to Video
High-quality image-to-video with Wan 2.5
Wan Video 2.5 Text to Video
High-quality text-to-video with Wan 2.5
Image models
44 pagesText-to-image, editing and vector output.
Seedream 5.0 Pro
ByteDance's flagship text-to-image model — deep-thinking prompt understanding, native text rendering in 14 languages, and precise control over dense layouts and structured designs
Midjourney V7
Midjourney V7 text-to-image generation with photorealistic and artistic outputs
GPT Image 2 Text to Image
OpenAI GPT Image 2 text-to-image. High-fidelity image generation with strong prompt adherence; supports up to 4K rendering.
Mai Image 2.5 Edit
Microsoft MAI-Image-2.5 image editing (pixel-level, cleanup, backgrounds, text)
Nano Banana 2
Ultra-fast lightweight image generation with Nano Banana 2 for rapid creative iteration
Grok Imagine Image Quality
xAI Grok Imagine Image (Quality tier) text-to-image. Quality-optimized generation up to 2K resolution with a wide range of aspect ratios.
HunyuanImage 3.0 Instruct Edit
HunyuanImage 3.0 Instruct Edit is an advanced instruction-following image editing model by Tencent. It allows users to modify and transform images based on natural language instructions while preserving the original content and structure of the image.
HiDream O1 Image
HiDream O1 pixel-native text-to-image that reasons before it draws. Unified model producing high-resolution images up to 2K, with optional subject references for personalization.
Luma UNI 1 Max
Luma Uni-1 (Max) — a reasoning-first image model that plans intent before generating, tuned for art-directed, cinematic results with optional reference images.
Flux 2 Max
Maximum quality image generation with the Flux 2 Max model for ultra-detailed photorealistic outputs
Kling Image 3.0
Kling V3 text-to-image generation with high-quality photorealistic and artistic outputs
Kling Image O1
Kling Image O1 is an advanced image generation model by Kuaishou that produces high-quality, detailed images from text prompts. It leverages state-of-the-art diffusion techniques to deliver visually compelling results for creative and professional use cases.
Seedream 4.0
ByteDance Seedream 4.0 text-to-image. High-resolution generation (up to 4K) with multi-image output and built-in prompt enhancement.
Wan 2.7 Pro Edit
WAN 2.7 Pro delivers professional-grade image editing and transformation with text-guided instructions, supporting up to 4 reference images for precise modifications
Flux 2 Flex
Flexible image generation and editing model from the Flux 2 family with versatile creative controls
HiDream O1 Image Edit
HiDream O1 image editing. Instruction-driven edits and subject-driven changes on a supplied image at up to 2K resolution.
Ideogram V4
Ideogram V4 text-to-image. Creates high-quality images, posters, and logos with crisp visuals, accurate text rendering, fine detail, and full creative control.
Krea 2 Medium
Krea 2 Medium text-to-image. Fast, photoreal generation with tunable creativity levels and optional style references.
Flux 2 Klein 9B Base
Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
Wan 2.6 Text to Image
Text-to-image generation with WAN V2.6
Qwen Image Edit 2511
Qwen Image Edit 2511 is an image editing model from Alibaba that enables precise modifications to existing images based on textual instructions. It supports a wide range of editing tasks including object removal, style transfer, and content transformation.
HiDream O1 Image Dev
HiDream O1 Image (Dev) — the faster, distilled tier of the pixel-native text-to-image model. Lower cost with strong prompt adherence up to 2K.
ERNIE Image Turbo
Baidu ERNIE Image (Turbo) — an 8B diffusion-transformer distilled for fast generation, strong at posters, comics, and text-in-image layouts.
ImagineArt 2.0
ImagineArt 2.0 text-to-image with true-to-life realism, precise text rendering, and cinematic lighting. Supports 1K/2K output and low/high reasoning.
Recraft 4.1 Text to Image
Recraft V4.1 standard text-to-image. High-quality image generation from prompts with strong style control and consistency.
Flux 2 Flash
Ultra-fast image generation with Flux 2
Recraft V4
Latest Recraft V4 text-to-image model with exceptional design quality and typography support
Flux 2 Klein 4B Base Edit
Flux 2 Klein 4B Base Edit is a compact yet powerful image editing model built on the Flux 2 architecture. It enables high-quality instruction-based image edits at an efficient parameter scale, making it suitable for fast and responsive image modification workflows.
Flux Kontext
Context-aware generation with coherent compositions
Imagen 4
Premium model for photorealistic results (text-only)
Qwen Image
High-quality realistic images with excellent detail
Flux 1.1 Pro
Professional grade Flux model with enhanced features
Qwen Z Image
Advanced image generation model with enhanced creative capabilities
Flux Schnell
Fast image generation model with high-quality outputs and efficient processing
Runway Gen4 Image
Runway Gen4 image generation model with high-quality creative outputs and style control
Cosmos 3 Super
NVIDIA Cosmos 3 Super text-to-image. High-fidelity generation with optional agentic refinement and prompt expansion for strong prompt adherence.
FireRed Image Edit
Advanced image editing model for prompt-guided modifications, inpainting, and creative edits
LongCat Edit
Longcat Edit is an image editing model that supports long-context image understanding and precise editing based on detailed textual instructions. It excels at complex multi-step edits and nuanced image transformations.
Midjourney Niji 6
Midjourney Niji 6 anime and illustration-focused image generation optimized for anime, manga, and cartoon styles
Nano Banana
Fast and efficient image generation with vibrant colors
Qwen Image 2 Edit
Qwen Image 2 Edit is an image editing model that allows users to modify and enhance images based on specific instructions. It is designed to provide high-quality edits while maintaining the integrity of the original image, making it suitable for photo editing, retouching, and creative transformations.
Recraft 4 Text to Image
Recraft 4 is a powerful text-to-image model that generates high-quality images from textual descriptions. It is designed to create detailed and visually appealing images based on user input, making it ideal for various applications such as content creation, design, and more.
SeedVR Upscale
VR-optimized image upscaling
Topaz Image Colorization
AI-powered colorization of black-and-white images with natural, realistic color restoration
Lipsync models
13 pagesTalking-head and avatar models that sync mouth to audio.
HeyGen Avatar V3
Generate digital twin videos with HeyGen Avatar V3. Choose from 700+ premade avatars with customizable voices, expressions, and styles.
HeyGen Avatar V5
HeyGen Avatar V digital twins — natural talking-avatar videos from a script or audio with premium lip-sync
HeyGen Image to Video
Turn your photo into a talking avatar video. Upload a selfie and choose a voice — HeyGen Avatar 4 generates realistic lip-sync video.
Sync React 1
Emotional lipsync with facial expressions and head movement. Supports emotions (happy, angry, sad, neutral, disgusted, surprised), model modes (lips, face, head), and temperature control.
VEED Avatars
Professional talking avatar generation with natural expressions and lip synchronization
Kling Avatar Pro
Professional quality talking avatar with advanced lip synchronization
LTX 2 Audio to Video
Generate videos from audio input with synchronized visuals and motion
LTX 2.3 Audio to Video
High-quality, fast AI video model for audio-to-video with lipsync, stylized transform capabilities
Sync Lipsync 2.0
Professional video lipsync with sync mode control. Supports cut_off, loop, bounce, silence, and remap sync modes.
VEED Fabric 1.0
Professional fabric and texture video generation
VEED Fabric 1.0 Text
Turn a photo into a talking video from just a script — the voice is auto-generated to match the character
VEED Lipsync
Fast and affordable video lipsync generation. Simple video + audio input.
Wan 2.2 Speech to Video
Generates high-quality videos from static images and audio with realistic facial expressions, body movements, and professional camera work
Audio & voice models
13 pagesSpeech, voice cloning and music generation.
Cartesia Sonic 3.5
Cartesia's newest and top-ranked TTS model (Sonic 3.5). Multilingual, highly expressive, with emotion control, speed tuning, and voice cloning.
Inworld TTS 1.5 Max
Inworld Realtime TTS 1.5 Max — the #1 ranked Inworld model, delivering the best balance of quality and speed. Expressive, multilingual (15 languages), with voice library support.
Inworld TTS 2
Inworld Realtime TTS-2 (Research Preview) — Inworld's most powerful and expressive model. 100+ languages with natural-language style steering for fine-grained delivery control.
Seed Audio 1.0
ByteDance Seed Audio 1.0 — high-quality, natural-sounding text-to-speech with preset voices, optional reference audio (@Audio1–@Audio3) or a reference image, and speed/pitch/volume controls.
Gemini 3.1 Flash TTS
Google Gemini 3.1 Flash TTS — expressive text-to-speech with 30 voices, natural-language style control, inline audio tags ([sigh], [laughing], [whispering]), multilingual synthesis, and multi-speaker dialogue.
ElevenLabs Multilingual
Multilingual text-to-speech powered by ElevenLabs Multilingual v2.
Qwen 3 TTS 0.6B
Compact text-to-speech model with natural voice synthesis and efficient processing
Cartesia Voice Clone
Clone any voice using Cartesia AI. Upload a short audio sample to instantly create a personalized voice for text-to-speech generation.
ElevenLabs Voice Change
Voice conversion model that transforms voice characteristics while preserving speech content
Grok TTS
xAI Grok TTS - high-quality text-to-speech with 5 expressive voices, 21 languages, and speech tags support. Up to 15,000 characters per request.
Inworld TTS
High-quality text-to-speech powered by Inworld AI. Supports a curated library of expressive voices with multilingual capability.
Qwen 3 TTS Voice Design
Qwen 3 text-to-speech with custom voice design
Suno Sounds V5.5
Suno Sounds V5.5 is the latest sound generation model with improved quality, looping, tempo, key controls, and lyrics subtitle capture
Video upscaling
1 pageResolution and detail enhancement for footage you already have.
Every model. One app.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.