Generation modes
What a model takes in and what it hands back — the shape of the job. 12 terms.
Audio-to-video
Audio-to-video generation drives the picture from a soundtrack: the audio is the primary input and the visuals are generated to agree with it.
Background removal
Background removal separates a subject from everything behind it and discards the rest, leaving a cutout you can place on any backdrop.
Image-to-image
Image-to-image takes a picture as its primary input and returns a changed picture — restyled, corrected or varied — instead of inventing one from nothing.
Image-to-video
Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.
Lipsync
Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.
Motion control
Motion control transfers the movement in a driving video onto a different subject — the performance stays, the performer changes.
Reference-to-video
Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.
Segmentation
Segmentation labels which pixels in each frame belong to a chosen object, producing a mask that other tools then act on.
Text-to-image
Text-to-image renders a still picture from a written description, with no picture going in.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
Video extend
Video extend continues an existing clip past its final frame, generating new footage that starts from where the old footage stopped.
Video-to-video
Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.
Generation controls
The inputs and switches you set before you press generate. 18 terms.
Aspect ratio
Aspect ratio is the proportion between an output's width and its height, written as two numbers — 16:9 for landscape, 9:16 for a full phone screen.
Camera control
Camera control is the set of inputs that decide how the virtual camera behaves during a generated clip — whether it pushes in, orbits, pans, or stays locked off.
CFG scale
CFG scale controls how strictly a model obeys your prompt, trading obedience against the model's own sense of what a natural image looks like.
Denoising strength
Denoising strength decides how much of your input picture gets thrown away before regeneration — low keeps it nearly intact, high keeps only the general shape.
Duration
Duration is how long a generated clip runs, chosen before generation from whatever lengths the model supports rather than trimmed afterwards.
First-last frame
First-last frame generation takes two stills — where the clip starts and where it ends — and generates the motion that gets from one to the other.
Frame rate
Frame rate is how many still frames make up one second of video, written as fps — 24 for a filmic look, 30 for broadcast-standard motion, 60 for smooth fast action.
Motion level
Motion level is a coarse dial for how much movement a generated clip contains, usually offered as low, medium or high rather than a precise number.
Negative prompt
A negative prompt lists what you do not want in the output, steering the generation away from those terms instead of towards them.
Prompt
A prompt is the written instruction a generative model reads to decide what to make — the one input almost every model requires.
Prompt expansion
Prompt expansion is a provider-side step that rewrites your short prompt into a longer, more detailed one before the generator ever sees it.
Reference image
A reference image is a picture supplied alongside the prompt so the model can copy an identity, product or style from it, without that picture becoming a frame of the output.
Resolution
Resolution is how many pixels an output contains, usually named by its height — 720p, 1080p, 4K — and set before generation rather than after.
Safety checker
A safety checker is an automated filter that inspects prompts, outputs or both and blocks material a provider does not permit.
Sampling steps
Sampling steps is how many passes a model takes to turn its starting noise into a finished output — more passes, more refinement, more time.
Seed
A seed is the number that decides the random starting noise for a generation, so the same seed with the same settings reproduces the same output.
Style preset
A style preset is a named look — cinematic, anime, documentary — that a model applies without you having to describe it in the prompt.
Watermark
A watermark is a mark applied to generated media to identify where it came from — either visible on the picture or embedded invisibly in the file.
Output quality
The words people reach for when an output is nearly right. 8 terms.
Character consistency
Character consistency is whether the same person, mascot or product still looks like itself across separate generations.
Frame interpolation
Frame interpolation invents new frames between existing ones, raising a clip's frame rate or slowing it down without it becoming choppy.
Inpainting
Inpainting regenerates a region you have masked while leaving the rest of the picture untouched, so a change stays local.
Native audio
Native audio means a video model generates its own soundtrack — dialogue, effects, ambience — in the same pass as the picture, rather than leaving you a silent clip.
Outpainting
Outpainting extends an image beyond its original borders, generating new content that continues the scene outward.
Prompt adherence
Prompt adherence is how faithfully a model does what the prompt actually said, as opposed to producing something attractive in the same neighbourhood.
Temporal consistency
Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.
Upscaling
Upscaling raises the resolution of media you already have, inventing plausible detail rather than recovering detail that was never recorded.
Models and architecture
How the models are built, adapted and ranked. 10 terms.
Autoregressive model
An autoregressive model generates one piece at a time, each piece conditioned on everything produced before it, rather than refining a whole output at once.
Diffusion model
A diffusion model generates by starting from random noise and removing a little of it at a time until a picture or clip is left behind.
Diffusion transformer
A diffusion transformer is a diffusion model whose internals are a transformer — the same architecture behind large language models — instead of the convolutional network earlier image models used.
Distillation
Distillation trains a smaller or faster model to imitate a larger one's outputs, which is where the fast and turbo variants of familiar models come from.
Elo rating
An Elo rating ranks models by head-to-head preference: people compare two outputs from the same prompt, and each model's number moves according to who won and how strong the opponent was.
Fine-tuning
Fine-tuning continues training an existing model on your own examples so it produces your subject or style by default, rather than on request.
Flow matching
Flow matching trains a model to follow a direct path from noise to data, rather than learning to reverse a long chain of noise-adding steps.
Latent space
Latent space is the compressed representation a model actually works in — a much smaller version of the image or clip that keeps meaning while discarding raw pixel count.
LoRA
A LoRA is a small add-on file that adjusts a large model's behaviour — teaching it a specific character, product or style — without retraining or replacing the model itself.
Multimodal model
A multimodal model handles more than one kind of data — text, images, audio, video — inside a single system rather than bolting separate tools together.
Speech, voice and audio
Voice, dubbing, transcription and music vocabulary. 12 terms.
AI dubbing
AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.
Audio tags
Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.
Forced alignment
Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.
Speech-to-speech
Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.
Speech-to-text
Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.
Stem separation
Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.
Text-to-music
Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
Voice design
Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.
Voice isolation
Voice isolation separates speech from everything else in a recording — traffic, room noise, music — and keeps only the voice.
Voice stability
Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.
Publishing and performance
What the people commissioning the work call it — what gets watched, what gets skipped, and what ships again next week. 10 terms.
B-roll
B-roll is the supporting footage cut over narration or an interview — everything on screen that is not the person doing the talking.
Batch generation
Batch generation runs one brief as several jobs at once — multiple variants, models or formats — so the comparison happens in a single sitting instead of across a week.
Brand kit
A brand kit is the saved set of colours, fonts, logo, product images, tone of voice and default frame shape that generations are expected to obey without being reminded.
Burned-in captions
Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.
Content repurposing
Repurposing is turning one finished piece into the several format-native cuts each platform expects, rather than posting the same file to all of them.
Content velocity
Content velocity is how much finished, publishable work actually ships per week — the throughput of the pipeline rather than the quality of any single piece.
Creative fatigue
Creative fatigue is performance decaying because the audience has seen the same piece too many times, rather than because anything about the piece changed.
Hook rate
Hook rate is the share of people shown a video who are still watching a few seconds in — the number that grades the opening, not the edit behind it.
Retention curve
A retention curve shows how many of the people who started a video are still watching at each second of it, and watch time is the area under that curve.
UGC ad
A UGC ad is an advert made to look like an ordinary person's own post — handheld, spoken to camera, unpolished on purpose — so it reads as a recommendation rather than a commercial.
Every term, A to Z
Looking for something else?
This is vocabulary, not instructions. For the settings a specific model exposes see the model catalog; for how to actually do a job see the editing tasks or the workflow recipes.