Glossary

    The AI video, image and audio glossary

    70 terms you will meet on a generation screen, defined in plain English — what each one does, what it costs you when it is wrong, and which models expose it.

    Generation modes

    What a model takes in and what it hands back — the shape of the job. 12 terms.

    Audio-to-video

    Audio-to-video generation drives the picture from a soundtrack: the audio is the primary input and the visuals are generated to agree with it.

    Background removal

    Background removal separates a subject from everything behind it and discards the rest, leaving a cutout you can place on any backdrop.

    Image-to-image

    Image-to-image takes a picture as its primary input and returns a changed picture — restyled, corrected or varied — instead of inventing one from nothing.

    Image-to-video

    Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.

    Lipsync

    Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.

    Motion control

    Motion control transfers the movement in a driving video onto a different subject — the performance stays, the performer changes.

    Reference-to-video

    Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.

    Segmentation

    Segmentation labels which pixels in each frame belong to a chosen object, producing a mask that other tools then act on.

    Text-to-image

    Text-to-image renders a still picture from a written description, with no picture going in.

    Text-to-video

    Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.

    Video extend

    Video extend continues an existing clip past its final frame, generating new footage that starts from where the old footage stopped.

    Video-to-video

    Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.

    Generation controls

    The inputs and switches you set before you press generate. 18 terms.

    Aspect ratio

    Aspect ratio is the proportion between an output's width and its height, written as two numbers — 16:9 for landscape, 9:16 for a full phone screen.

    Camera control

    Camera control is the set of inputs that decide how the virtual camera behaves during a generated clip — whether it pushes in, orbits, pans, or stays locked off.

    CFG scale

    CFG scale controls how strictly a model obeys your prompt, trading obedience against the model's own sense of what a natural image looks like.

    Denoising strength

    Denoising strength decides how much of your input picture gets thrown away before regeneration — low keeps it nearly intact, high keeps only the general shape.

    Duration

    Duration is how long a generated clip runs, chosen before generation from whatever lengths the model supports rather than trimmed afterwards.

    First-last frame

    First-last frame generation takes two stills — where the clip starts and where it ends — and generates the motion that gets from one to the other.

    Frame rate

    Frame rate is how many still frames make up one second of video, written as fps — 24 for a filmic look, 30 for broadcast-standard motion, 60 for smooth fast action.

    Motion level

    Motion level is a coarse dial for how much movement a generated clip contains, usually offered as low, medium or high rather than a precise number.

    Negative prompt

    A negative prompt lists what you do not want in the output, steering the generation away from those terms instead of towards them.

    Prompt

    A prompt is the written instruction a generative model reads to decide what to make — the one input almost every model requires.

    Prompt expansion

    Prompt expansion is a provider-side step that rewrites your short prompt into a longer, more detailed one before the generator ever sees it.

    Reference image

    A reference image is a picture supplied alongside the prompt so the model can copy an identity, product or style from it, without that picture becoming a frame of the output.

    Resolution

    Resolution is how many pixels an output contains, usually named by its height — 720p, 1080p, 4K — and set before generation rather than after.

    Safety checker

    A safety checker is an automated filter that inspects prompts, outputs or both and blocks material a provider does not permit.

    Sampling steps

    Sampling steps is how many passes a model takes to turn its starting noise into a finished output — more passes, more refinement, more time.

    Seed

    A seed is the number that decides the random starting noise for a generation, so the same seed with the same settings reproduces the same output.

    Style preset

    A style preset is a named look — cinematic, anime, documentary — that a model applies without you having to describe it in the prompt.

    Watermark

    A watermark is a mark applied to generated media to identify where it came from — either visible on the picture or embedded invisibly in the file.

    Output quality

    The words people reach for when an output is nearly right. 8 terms.

    Models and architecture

    How the models are built, adapted and ranked. 10 terms.

    Autoregressive model

    An autoregressive model generates one piece at a time, each piece conditioned on everything produced before it, rather than refining a whole output at once.

    Diffusion model

    A diffusion model generates by starting from random noise and removing a little of it at a time until a picture or clip is left behind.

    Diffusion transformer

    A diffusion transformer is a diffusion model whose internals are a transformer — the same architecture behind large language models — instead of the convolutional network earlier image models used.

    Distillation

    Distillation trains a smaller or faster model to imitate a larger one's outputs, which is where the fast and turbo variants of familiar models come from.

    Elo rating

    An Elo rating ranks models by head-to-head preference: people compare two outputs from the same prompt, and each model's number moves according to who won and how strong the opponent was.

    Fine-tuning

    Fine-tuning continues training an existing model on your own examples so it produces your subject or style by default, rather than on request.

    Flow matching

    Flow matching trains a model to follow a direct path from noise to data, rather than learning to reverse a long chain of noise-adding steps.

    Latent space

    Latent space is the compressed representation a model actually works in — a much smaller version of the image or clip that keeps meaning while discarding raw pixel count.

    LoRA

    A LoRA is a small add-on file that adjusts a large model's behaviour — teaching it a specific character, product or style — without retraining or replacing the model itself.

    Multimodal model

    A multimodal model handles more than one kind of data — text, images, audio, video — inside a single system rather than bolting separate tools together.

    Speech, voice and audio

    Voice, dubbing, transcription and music vocabulary. 12 terms.

    AI dubbing

    AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.

    Audio tags

    Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.

    Forced alignment

    Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.

    Speech-to-speech

    Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.

    Speech-to-text

    Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.

    Stem separation

    Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.

    Text-to-music

    Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.

    Text-to-speech

    Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.

    Voice cloning

    Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.

    Voice design

    Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.

    Voice isolation

    Voice isolation separates speech from everything else in a recording — traffic, room noise, music — and keeps only the voice.

    Voice stability

    Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.

    Publishing and performance

    What the people commissioning the work call it — what gets watched, what gets skipped, and what ships again next week. 10 terms.

    B-roll

    B-roll is the supporting footage cut over narration or an interview — everything on screen that is not the person doing the talking.

    Batch generation

    Batch generation runs one brief as several jobs at once — multiple variants, models or formats — so the comparison happens in a single sitting instead of across a week.

    Brand kit

    A brand kit is the saved set of colours, fonts, logo, product images, tone of voice and default frame shape that generations are expected to obey without being reminded.

    Burned-in captions

    Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.

    Content repurposing

    Repurposing is turning one finished piece into the several format-native cuts each platform expects, rather than posting the same file to all of them.

    Content velocity

    Content velocity is how much finished, publishable work actually ships per week — the throughput of the pipeline rather than the quality of any single piece.

    Creative fatigue

    Creative fatigue is performance decaying because the audience has seen the same piece too many times, rather than because anything about the piece changed.

    Hook rate

    Hook rate is the share of people shown a video who are still watching a few seconds in — the number that grades the opening, not the edit behind it.

    Retention curve

    A retention curve shows how many of the people who started a video are still watching at each second of it, and watch time is the area under that curve.

    UGC ad

    A UGC ad is an advert made to look like an ordinary person's own post — handheld, spoken to camera, unpolished on purpose — so it reads as a recommendation rather than a commercial.

    Every term, A to Z

    Looking for something else?

    This is vocabulary, not instructions. For the settings a specific model exposes see the model catalog; for how to actually do a job see the editing tasks or the workflow recipes.