Glossary

    The AI video, image and audio glossary

    232 terms you will meet on a generation screen, defined in plain English — what each one does, what it costs you when it is wrong, and which models expose it.

    Generation modes

    What a model takes in and what it hands back — the shape of the job. 28 terms.

    Audio-to-video

    Audio-to-video generation drives the picture from a soundtrack: the audio is the primary input and the visuals are generated to agree with it.

    Background removal

    Background removal separates a subject from everything behind it and discards the rest, leaving a cutout you can place on any backdrop.

    Beat list

    A beat list is the shot list written before a 30-second pass: each beat is one generate, because one clip will not carry the whole story.

    Digital-twin creator

    A digital-twin creator is a persistent likeness of a real person — face, voice, manner — used as a reusable presenter, not a stock avatar from a roster.

    Driving clip

    A driving clip is the video that supplies pose and timing to a motion-control generate — the performance a still or character then wears.

    Extend retake

    An extend-retake regenerates one continuation or a marked span of a clip while keeping the rest, rather than re-rolling the whole generate.

    Image-to-image

    Image-to-image takes a picture as its primary input and returns a changed picture — restyled, corrected or varied — instead of inventing one from nothing.

    Image-to-video

    Image-to-video animates a still you supply: the picture becomes the opening frame, and the prompt describes only what happens next.

    Lipsync

    Lipsync generation drives a face's mouth from an audio track, so the speech reads as spoken rather than dubbed over the top.

    Motion control

    Motion control transfers the movement in a driving video onto a different subject — the performance stays, the performer changes.

    Multi-reference

    Multi-reference generation conditions a clip on several stills at once — a stack of identities or products, not one picture and a prompt.

    Multi-scene generation

    Multi-scene generation is building a video as separately generated shots that are then stitched, rather than asking one clip to carry the whole story.

    One-tap template

    A one-tap template is a curated, photo-in video format you run as a finished clip — no scene plan — rather than a multi-scene workflow you author.

    Picture-in-picture

    Picture-in-picture is compositing a smaller video — usually a talking head or reaction — over a base clip, rather than generating a new clip that already contains both.

    Premade avatar

    A premade avatar is a stock presenter from a ready-made roster you pick, rather than a face you supply or a mouth you lipsync onto footage you already shot.

    Reference stack

    A reference stack is the set of stills and clips a reference-to-video job consumes — identities the model may place, not frame one.

    Reference-to-video

    Reference-to-video builds a clip around subjects supplied as separate reference images, rather than starting from one fixed opening frame.

    Region edit

    A region edit changes only a boxed area of a clip across frames — swap a label, fix a hand — while the rest of the shot is preserved.

    Reusable character

    A reusable character is a named workflow asset — reference images saved once under a key — that later scenes in that recipe call without re-uploading.

    Segmentation

    Segmentation labels which pixels in each frame belong to a chosen object, producing a mask that other tools then act on.

    Slideshow

    A slideshow is a sequence of stills — generated or uploaded — assembled as a carousel or a short video, rather than a single generated clip.

    Still contract

    A still contract is the rule that an image-to-video clip must obey its input still — identity, product, framing — rather than drifting into a new picture.

    Storyboard

    A storyboard is a planned sequence of shots — stills, prompts, or both — that a later generate follows, rather than one prompt asked to carry the whole film.

    Text overlay

    Text overlay is copy you authored and burned onto a clip, not a speech transcript and not letters a model painted into the scene.

    Text-to-image

    Text-to-image renders a still picture from a written description, with no picture going in.

    Text-to-video

    Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.

    Video extend

    Video extend continues an existing clip past its final frame, generating new footage that starts from where the old footage stopped.

    Video-to-video

    Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.

    Generation controls

    The inputs and switches you set before you press generate. 27 terms.

    9:16 vs 4:5 vs 1:1

    9:16 vs 4:5 vs 1:1 is the short-form frame choice: full-screen Stories and Reels, tall feed posts, or square tiles — three compositions, not three crops.

    Aspect ratio

    Aspect ratio is the proportion between an output's width and its height, written as two numbers — 16:9 for landscape, 9:16 for a full phone screen.

    Camera control

    Camera control is the set of inputs that decide how the virtual camera behaves during a generated clip — whether it pushes in, orbits, pans, or stays locked off.

    Camera path

    A camera path is a planned trajectory for the virtual camera — keyframed push, orbit, or crane — not a single motion word and not a motion-level dial.

    CFG scale

    CFG scale controls how strictly a model obeys your prompt, trading obedience against the model's own sense of what a natural image looks like.

    Denoising strength

    Denoising strength decides how much of your input picture gets thrown away before regeneration — low keeps it nearly intact, high keeps only the general shape.

    Duration

    Duration is how long a generated clip runs, chosen before generation from whatever lengths the model supports rather than trimmed afterwards.

    First-last frame

    First-last frame generation takes two stills — where the clip starts and where it ends — and generates the motion that gets from one to the other.

    Frame rate

    Frame rate is how many still frames make up one second of video, written as fps — 24 for a filmic look, 30 for broadcast-standard motion, 60 for smooth fast action.

    Generation time

    Generation time is how long a job takes to finish once you submit it — driven by resolution, step count and model size — and it is routinely confused with duration, which is how long the finished clip plays.

    Motion level

    Motion level is a coarse dial for how much movement a generated clip contains, usually offered as low, medium or high rather than a precise number.

    Native resolution

    Native resolution is the pixel size a catalog row actually generates — the listed output — not an upscale of a smaller pass sold as the same label.

    Negative prompt

    A negative prompt lists what you do not want in the output, steering the generation away from those terms instead of towards them.

    Negative prompt (video)

    A video negative prompt lists motion and time failures to avoid — flicker, morphing, extra limbs, camera drift — not just still-image artefacts.

    Prompt

    A prompt is the written instruction a generative model reads to decide what to make — the one input almost every model requires.

    Prompt expansion

    Prompt expansion is a provider-side step that rewrites your short prompt into a longer, more detailed one before the generator ever sees it.

    Prompt weighting

    Prompt weighting is explicit syntax inside a prompt — brackets, parentheses or a colon-and-number — that turns one word or phrase up or down without changing the strength of everything else in the prompt.

    Reference image

    A reference image is a picture supplied alongside the prompt so the model can copy an identity, product or style from it, without that picture becoming a frame of the output.

    Resolution

    Resolution is how many pixels an output contains, usually named by its height — 720p, 1080p, 4K — and set before generation rather than after.

    Safety checker

    A safety checker is an automated filter that inspects prompts, outputs or both and blocks material a provider does not permit.

    Sampling steps

    Sampling steps is how many passes a model takes to turn its starting noise into a finished output — more passes, more refinement, more time.

    Seed

    A seed is the number that decides the random starting noise for a generation, so the same seed with the same settings reproduces the same output.

    Seed lock

    A seed-lock freezes the random seed across runs so prompt or setting changes are the only variable — the honest way to iterate a nearly-right generate.

    Style preset

    A style preset is a named look — cinematic, anime, documentary — that a model applies without you having to describe it in the prompt.

    Video input

    Video input is a catalog constraint: the model will not run unless you upload footage, as distinct from a still-in job or a text-only prompt.

    Visible watermark

    A visible watermark is a logo or wordmark drawn on the picture — optional chrome a crop can remove, unlike an invisible provenance mark.

    Watermark

    A watermark is a mark applied to generated media to identify where it came from — either visible on the picture or embedded invisibly in the file.

    Credits and billing

    What a generation actually costs to run and how that cost is metered — the credit itself, and the mechanism that turns a job into a number of them. Not what Versely charges for a plan; /cost owns that. 5 terms.

    Output quality

    The words people reach for when an output is nearly right. 15 terms.

    4K native

    4K native means the model generated at 3840×2160 (or a 4K tier), not that a 720p or 1080p clip was upscaled afterwards.

    Alpha channel

    An alpha channel is a fourth layer of image data, alongside red, green and blue, that records how transparent each pixel is — and whether a file keeps one depends entirely on its format, not on how good the cutout underneath it was.

    Bit depth

    Bit depth is how many distinct brightness levels a file can record per colour channel — 256 at 8-bit, 1,024 at 10-bit — and it is the axis responsible for banding, a defect that more resolution or a sharper export setting cannot fix because neither one touches it.

    Character consistency

    Character consistency is whether the same person, mascot or product still looks like itself across separate generations.

    Frame interpolation

    Frame interpolation invents new frames between existing ones, raising a clip's frame rate or slowing it down without it becoming choppy.

    Inpainting

    Inpainting regenerates a region you have masked while leaving the rest of the picture untouched, so a change stays local.

    Native audio

    Native audio means a video model generates its own soundtrack — dialogue, effects, ambience — in the same pass as the picture, rather than leaving you a silent clip.

    Outpainting

    Outpainting extends an image beyond its original borders, generating new content that continues the scene outward.

    Physics consistency

    Physics consistency is whether generated motion obeys plausible weight, contact and inertia — a glass that falls, a cloth that hangs, a body that has mass.

    Prompt adherence

    Prompt adherence is how faithfully a model does what the prompt actually said, as opposed to producing something attractive in the same neighbourhood.

    Scene continuity

    Scene continuity is whether identity, wardrobe, set and lighting hold across separately generated shots, rather than inside one clip's temporal consistency.

    Silent row

    A silent row is a catalog video model whose audio field is false or null — it returns picture only, so speech is a later pass.

    Temporal consistency

    Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.

    Upscaling

    Upscaling raises the resolution of media you already have, inventing plausible detail rather than recovering detail that was never recorded.

    Video inpainting

    Video inpainting regenerates a masked region through time, so a removal or replacement holds across frames instead of flickering per still.

    Models and architecture

    How the models are built, adapted and ranked. 12 terms.

    Any-to-any

    Any-to-any is a model that takes mixed inputs — text, image, audio, video — and returns mixed outputs, rather than a one-way image-to-video pipe.

    Autoregressive model

    An autoregressive model generates one piece at a time, each piece conditioned on everything produced before it, rather than refining a whole output at once.

    Diffusion model

    A diffusion model generates by starting from random noise and removing a little of it at a time until a picture or clip is left behind.

    Diffusion transformer

    A diffusion transformer is a diffusion model whose internals are a transformer — the same architecture behind large language models — instead of the convolutional network earlier image models used.

    Distillation

    Distillation trains a smaller or faster model to imitate a larger one's outputs, which is where the fast and turbo variants of familiar models come from.

    Elo rating

    An Elo rating ranks models by head-to-head preference: people compare two outputs from the same prompt, and each model's number moves according to who won and how strong the opponent was.

    Fine-tuning

    Fine-tuning continues training an existing model on your own examples so it produces your subject or style by default, rather than on request.

    Flow matching

    Flow matching trains a model to follow a direct path from noise to data, rather than learning to reverse a long chain of noise-adding steps.

    Latent space

    Latent space is the compressed representation a model actually works in — a much smaller version of the image or clip that keeps meaning while discarding raw pixel count.

    LoRA

    A LoRA is a small add-on file that adjusts a large model's behaviour — teaching it a specific character, product or style — without retraining or replacing the model itself.

    Multimodal model

    A multimodal model handles more than one kind of data — text, images, audio, video — inside a single system rather than bolting separate tools together.

    SynthID

    SynthID is Google DeepMind's invisible watermark for AI media — a detector-readable signal in the output, not a logo in the corner.

    Speech, voice and audio

    Voice, dubbing, transcription and music vocabulary. 23 terms.

    AI dubbing

    AI dubbing replaces a video's spoken audio with another language, usually keeping the original speaker's voice and optionally re-syncing their mouth.

    ASR

    ASR (automatic speech recognition) is the engine that turns speech into text — the pass behind captions, transcripts and word timestamps.

    Audio extend

    Audio extend continues a previously generated music track from a chosen point, rather than writing a new song or time-stretching the file.

    Audio tags

    Audio tags are markers written inside the text of a script — bracketed or angle-bracketed cues like a laugh or a whisper — that tell a speech model how to deliver the words around them.

    Cover song

    A cover song, in generation, is a new arrangement of an existing track you point at, not an original composition from a blank prompt.

    Diegetic sound

    Diegetic sound belongs in the scene — a door slam, a voice in the room, a radio on set — as opposed to score, voiceover or a caption of that sound.

    Forced alignment

    Forced alignment matches a known transcript to the audio it came from, working out exactly when each word was spoken.

    Loudness

    Loudness is a measured, standardised value — expressed in LUFS — that predicts how loud a track will sound to a listener over its whole length, as distinct from a peak level, which only ever measures the single highest sample.

    Multi-speaker dialogue

    Multi-speaker dialogue is one speech generation with a distinct voice per speaker alias, so a conversation is a single take instead of several stitched reads.

    Off-screen speaker

    An off-screen speaker is speech with no mouth in the frame — a caption speaker-ID and mix job, not a reason to generate a talking head.

    Phoneme-level lipsync

    Phoneme-level lipsync drives the mouth from individual speech sounds, not from a rough viseme or a per-word flap, so consonants actually land.

    Speech-to-speech

    Speech-to-speech takes a recording of one person talking and re-renders it in a different voice, keeping the original performance intact.

    Speech-to-text

    Speech-to-text converts spoken audio into written text, producing the transcript that captions, translation and search all depend on.

    Stem separation

    Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.

    Text-to-music

    Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.

    Text-to-SFX

    Text-to-SFX generates a short sound effect — a whoosh, hit, footstep or ambience — from a written description, rather than composing a song.

    Text-to-speech

    Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.

    Voice cloning

    Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.

    Voice design

    Voice design creates a new synthetic voice from a written description — age, accent, texture, energy — instead of cloning one from a recording.

    Voice isolation

    Voice isolation meaning: separate speech from traffic, room noise, and music so only the voice remains. Not stem separation.

    Voice stability

    Voice stability is the control that decides how much a synthetic voice varies its delivery — steady and predictable at one end, expressive and unpredictable at the other.

    Voiceover

    A voiceover is a spoken track laid under picture, rather than lipsync that drives a mouth or a caption track that visualises speech.

    Word timestamp

    A word timestamp is the start (and usually end) time of one spoken token — the clock karaoke captions and word-highlight styles consume.

    Publishing and performance

    What the people commissioning the work call it — what gets watched, what gets skipped, and what ships again next week. 69 terms.

    A-roll

    A-roll is the primary footage a cut is built around — usually a person speaking — as distinct from B-roll that covers it and from a generate that tries to be both.

    Agent memory

    Agent memory is a lasting preference or fact the agent stores across conversations, separate from the structured brand-kit fields it applies to a render.

    Average view duration

    Average view duration is mean watch time per view — total minutes divided by views — a typical stay, not a completion rate and not a total.

    B-roll

    B-roll is the supporting footage cut over narration or an interview — everything on screen that is not the person doing the talking.

    Batch generation

    Batch generation runs one brief as several jobs at once — multiple variants, models or formats — so the comparison happens in a single sitting instead of across a week.

    Brand kit

    A brand kit is the saved set of colours, fonts, logo, product images, tone of voice and default frame shape that generations are expected to obey without being reminded.

    Burned-in captions

    Burned-in captions are subtitles rendered into the video's pixels, so they cannot be switched off, restyled by the player, or lost when the file is re-uploaded somewhere else.

    Caption line length

    Caption line length is how many characters sit on one cue. Latin-script lines are commonly kept near 42 characters so a phone can still read them.

    Caption timing

    Caption timing is when each line appears and disappears relative to the speech — the clock, not the words and not the look.

    Characters per second

    Characters per second (CPS) is caption reading speed counted in characters: letters, spaces and punctuation shown per second of on-screen time.

    Chroma subsampling

    Chroma subsampling is how much colour detail a codec keeps relative to brightness detail, written as a three-number ratio like 4:2:0 or 4:4:4 — a separate axis from resolution, so a sharp, high-resolution file can still carry far less colour information than its pixel count implies.

    Closed captions

    Closed captions are a separate subtitle track the player draws, which the viewer can switch off and which each platform styles itself.

    Codec

    A codec is the scheme that compresses video or audio down to a storable size, and it is a different thing from a container format like MP4 or MOV, which is only the wrapper the codec's data travels inside.

    Cold open

    A cold open starts on the action — no logo, no greeting, no 'hey guys' — so the first frame is already the reason to stay.

    Comment bait

    Comment bait is an opening that asks for a comment as the hook — a reply prompt in frame one, graded by hook rate like any other open.

    Content credentials

    Content credentials are a cryptographically signed record of who or what made a file and what was done to it, attached by a hash computed over its exact bytes — which is why the record breaks the moment those bytes change, rather than degrading gracefully.

    Content repurposing

    Repurposing is turning one finished piece into the several format-native cuts each platform expects, rather than posting the same file to all of them.

    Content velocity

    Content velocity is how much finished, publishable work actually ships per week — the throughput of the pipeline rather than the quality of any single piece.

    Creative fatigue

    Creative fatigue is performance decaying because the audience has seen the same piece too many times, rather than because anything about the piece changed.

    CTA card

    A CTA card is an on-screen prompt to act — click, shop, follow, swipe — shown during the clip, unlike an end screen that only appears at the finish.

    End screen

    An end screen is the last-seconds overlay that points to another video, a subscribe, or a URL — YouTube's named object, not a burned-in caption.

    Engaged views

    Engaged views are YouTube's stay metric after 24 August 2026: a watch past the first frame, or a tap to watch, not a public view at playback start.

    Faceless channel

    A faceless channel publishes without showing the creator's face — voiceover, screen, b-roll, or a generated presenter instead of a to-camera host.

    First-frame view

    A first-frame view is YouTube's public view since 24 August 2026: the counter ticks when playback starts, from the first frame, not when someone stays.

    First-second hook

    A first-second hook is the opening beat that stops the thumb — frame one plus the first line — graded by hook rate, not by how the rest of the clip plays.

    Forced narrative

    Forced narrative captions only translate what a same-language viewer still cannot read — a sign, a letter, foreign speech — not the whole dialogue.

    Hold rate

    Hold rate is the share of people still watching after the hook window — the number that grades the body, once the opening has already done its job.

    Home feed

    A home feed is the algorithmically ranked For You / Reels / Home surface — what the platform chose to show, not a following timeline.

    Home Timeline impressions

    Home Timeline impressions are X (Twitter) counts of a post shown in the Home feed — not profile, Search, or replies — and the ones that can pay.

    Hook pack

    A hook pack is several distinct openings rendered against one brief and one body, so the variable under test is the first second.

    Hook rate

    Hook rate is the share of people shown a video who are still watching a few seconds in — the number that grades the opening, not the edit behind it.

    Hook–body–CTA

    Hook–body–CTA is the three-beat short-form shape: an opening that stops the thumb, a middle that pays it off, and a last line that asks for one action.

    Idea Pin

    An Idea Pin is Pinterest's 9:16 video surface — the name that still names that full-screen watch, even after the format was folded into Pins.

    Jump cut

    A jump cut is a hard edit between two similar frames of the same shot that skips time — a pacing tool in short-form, and sometimes an AI duration artefact.

    Karaoke captions

    Karaoke captions highlight each spoken word as it lands, using per-word timestamps instead of dropping a whole line on screen at once.

    Mute autoplay

    Mute autoplay is how most feeds start: the clip plays with sound off until the viewer unmutes, so picture and burned-in text have to carry the hook.

    Mute-safe

    Mute-safe creative is designed to work with sound off: picture and burned-in text carry the claim before anyone unmutes.

    Open captions

    Open captions are always-on subtitles in the picture: the broadcast name for burned-in captions, which the viewer cannot switch off.

    Originality score

    Originality score is a platform's internal grade of how much of a post is yours. A low score can zero out Rewards even when public views look fine.

    Overlay chrome

    Overlay chrome is the platform UI drawn on a 9:16 file — username, buttons, caption strip, AI chip — covering whatever you put there.

    Pack shot

    A pack shot is the product still that has to stay readable — label, shape, colour — the plate image-to-video is not allowed to melt.

    Pattern interrupt

    A pattern interrupt is a sudden change in picture, sound, or on-screen type that resets attention before the viewer swipes or the retention curve sags.

    Platform-native captions

    Platform-native captions are subtitles the player draws — YouTube CC, TikTok auto captions — not text burned into the file's pixels.

    Provenance

    Provenance is the recorded origin of a file — who or what made it and what was done to it — carried as a watermark, a signed manifest, or both.

    Qualified watch hours

    Qualified watch hours are YouTube Partner Program watch time that counts toward monetization — not every hour a public view counter could imply.

    Reading speed

    Caption reading speed is how fast on-screen words arrive, in words per minute. A cue gone before it can be read is a timing error, not a style one.

    Recurring series

    A recurring series is a saved workflow or slideshow premise put on a schedule, so each run writes a new episode against the same cast and format.

    Repost limit

    A repost limit is how often a platform will let the same file, or a near copy, circulate before it throttles reach or drops Rewards.

    Repurposing

    Repurposing is turning one source — a shoot, a webinar, a generate — into many assets, not posting the same file to every app.

    Retention curve

    A retention curve shows how many of the people who started a video are still watching at each second of it, and watch time is the area under that curve.

    Rewatch

    A rewatch is a second (or later) play of the same clip — loops, replays, and the rise on a retention curve past 100%.

    Safe area

    A safe area is the part of a vertical frame a platform's own interface will not cover — profile name, caption line, like and share buttons all sit outside it — and text or a face placed inside it is routinely hidden from the viewer entirely.

    Safe-title

    Safe-title is the broadcast inner frame that keeps type off TV overscan; social safe-area is a different problem — platform buttons, not a CRT crop.

    Saved workflow

    A saved workflow is a reusable video recipe — scenes, assets, style and model stored so you can run it again with a fresh plot.

    Saves

    Saves are bookmarks of a post — Instagram saves, TikTok favourites — a ranking signal that the viewer wants the file again, not just a like.

    Scheduler strip

    Scheduler strip is C2PA dying on a cross-post or API upload: the signed manifest is bound to one file's bytes, and the re-encode is a new file.

    SDH

    SDH is a caption track that includes speech and non-speech audio — [door slams], [music], speaker IDs — so the file is usable without sound.

    Shares per view

    Shares-per-view is shares divided by views: how often a play produces a send, and the usual proxy for whether a clip can travel past the original audience.

    Speaker ID

    A speaker ID is the caption label that names who is talking — [Maya], Host: — so a mute viewer can tell overlapping or off-screen voices apart.

    SRT

    SRT is the simplest caption sidecar: numbered cues, start and end times with a comma decimal, and plain text. No styling, no positioning.

    Swipe-away

    A swipe-away is a short-form skip: the viewer leaves before the clip holds. After YouTube's first-frame rule it can still write a public view.

    Syndication

    Syndication is publishing one piece on more than one surface — adapted per platform, not the same file fired at three apps.

    Talking head

    A talking head is a person speaking to camera — the A-roll spine of a clip — whether that person was filmed, cloned, or generated.

    Thumbnail

    A thumbnail is a still chosen or generated to represent a video in a feed, not the video itself and not a 16:9 generate that happens to look like a title card.

    Two-line limit

    The two-line limit is the captioning habit of keeping a cue to two rows so type stays readable and clear of platform chrome on a vertical frame.

    UGC ad

    A UGC ad is an advert made to look like an ordinary person's own post — handheld, spoken to camera, unpolished on purpose — so it reads as a recommendation rather than a commercial.

    Watch time

    Watch time is the total minutes people spent on a video — the area under the retention curve — not the public view count and not the completion rate.

    WebVTT

    WebVTT is the web-native caption file: a .vtt sidecar with cues, timestamps, and optional positioning and styling the player may honour.

    Word-highlight captions

    Word-highlight captions are burned-in subtitles that light up each word as it is spoken, using word-level timing and a caption-style accent colour.

    Marketing and creator economy

    The vocabulary of ad accounts, creator deals and disclosure law — spark ads and whitelisting, UGC and gifting, what the FTC and the EU actually require you to say. 53 terms.

    AI content label

    An AI content label is the platform-applied or creator-applied tag marking a post as made or meaningfully altered with AI — a visible badge, distinct from the legal disclosure obligations that exist independently of any single platform's UI.

    AI slop

    AI slop is unoriginal, mass-produced generative output — YouTube's July 2026 inauthentic buckets name the pattern, not a ban on using a model.

    AI UGC

    AI UGC is brand-owned, feed-native video generated to look like a person talking to camera about a product — UGC-style, not an unprompted customer post.

    AI-generated profile

    An AI-generated profile is Instagram's 31 August 2026 account label for a profile whose featured person is AI-made, not a human using AI tools.

    Article 50

    Article 50 of the EU AI Act is the 2 August 2026 transparency rule: labels on certain AI interactions, synthetic media, and public-interest AI text.

    Boosted post

    A boosted post takes something already published on a brand's own page and pays to extend its reach beyond who saw it organically, rather than building a new ad from scratch.

    Brand safety

    Brand safety is keeping content — and where it runs — away from anything that would embarrass a brand or misrepresent what it sells, and with generated content, the risk increasingly starts inside the asset itself rather than only in where it's placed.

    Cadence

    Cadence is the publishing rhythm you actually keep — three Reels a week, a daily Story — not velocity (how much you make) and not the calendar (which day).

    Completion rate

    Completion rate is the share of people who watched a video all the way to its last frame — one number, unlike the full retention curve, and the one most short-form platforms weight heavily in distribution.

    Conspicuous disclosure

    Conspicuous disclosure is New York's 'difficult to miss' bar for synthetic-performer ads: the notice has to be in the ad, not buried in a caption pile.

    Content calendar

    A content calendar is the dated plan of what publishes where — the schedule, not the render queue and not the velocity metric.

    Content Rewards

    Content Rewards are platform payout programmes that pay on qualified views or Home Timeline impressions of original posts — not on every public play.

    CPM vs. CPV

    CPM prices an ad by impressions delivered — cost per thousand shown — while CPV prices it by counted views, so the two buying models charge for fundamentally different things regardless of what either happens to cost on a given platform.

    Creativity Program

    TikTok's Creativity Program pays eligible original videos. Labelled AI stays in; undisclosed AI can mean removal. The rate does not change with the label.

    Dark posting

    A dark post is a paid social ad that is never published to a brand's own page or feed — it exists only as an ad, shown solely to the audience it's targeted at.

    Deepfake

    A deepfake, in Article 50, is AI audio, image or video that resembles real people, places or events and would appear authentic to a person.

    Demand gen

    Demand gen is YouTube and Meta's demand-generation placements — in-feed, Shorts or Reels, Discover-class surfaces — not skippable in-stream as such.

    Deployer

    A deployer, under the EU AI Act, is whoever uses an AI system under their authority. If you publish a deepfake, you owe the label the viewer sees.

    Duet

    A Duet is TikTok's side-by-side reaction format — a new video plays next to the original, both visible and both playing at once, letting a brand or creator respond to someone else's post without re-uploading it.

    Evergreen vs trend

    Evergreen vs trend is the split between clips that keep working for months and clips that only work while a sound, format, or news cycle is hot.

    Exclusivity window

    An exclusivity window is a contract clause stopping a creator from working with a named category of competing brands for a defined period — separate from, and priced separately from, the rights to use the content itself.

    Expressive-work exemption

    An expressive-work exemption lets NY skip ads for film, TV or games, while Article 50 still wants a non-spoiling label on artistic or satirical work.

    FTC endorsement disclosure

    FTC endorsement disclosure is the requirement that anyone endorsing a product with a material connection to the brand say so clearly and conspicuously, in a place and language a viewer can't miss — a rule that applies to a human creator, an AI avatar, or a cloned voice equally.

    Gifting campaign

    A gifting campaign sends free product to creators with no cash fee and usually no contractual guarantee of a post, trading the product itself for a chance at organic coverage.

    Inauthentic content

    Inauthentic content is YouTube's policy against mass-produced, repetitive, templated uploads — the disqualifier for YPP hours, not a ban on AI as such.

    Material connection

    A material connection is any relationship between an endorser and a brand that could affect how much weight a reasonable viewer gives the endorsement — payment, free or discounted product, employment, or a family or personal tie — and its existence, not its size, is what triggers a disclosure obligation.

    Micro-influencer

    A micro-influencer is a creator in the tier between nano (a small, personal following) and macro (celebrity-scale reach) — roughly tens of thousands of followers, with engagement and perceived authenticity that raw reach at the top of the ladder tends to lose.

    Native advertising

    Native advertising is paid content built to match the form and function of the platform it runs on — an in-feed video, a sponsored article — so it reads as a piece of that platform rather than an interruption to it.

    Overlay label

    An overlay label is an AI disclosure drawn on the video — YouTube Shorts' May 2026 placement — not a line under the player or in the description.

    Photoreal disclosure

    Photoreal disclosure is the duty to label AI media a viewer could mistake for a real person, place or event — realism, not whether a model was used.

    Product feed

    A product feed is the catalog of SKUs an ads system uses to assemble shopping creative — titles, images, prices — not a pack-shot still.

    Provider (AI Act)

    A provider under the EU AI Act places a generative system on the market and owes machine-readable marking of synthetic output, not the on-screen label.

    RPM

    RPM is revenue per mille: what a creator actually earns per thousand views or impressions, after the platform's cut, not what advertisers paid (CPM).

    Seeding

    Seeding is placing finished content into communities — subreddits, Discord servers, group chats — as a genuine contribution rather than an ad, so it earns distribution before or without any paid push.

    Share of voice

    Share of voice is a brand's portion of the total conversation in its category — mentions, tags and citations relative to named competitors — tracked over time as a competitive gauge rather than a per-post metric.

    Sound-on strategy

    Sound-on strategy is designing a video to reward viewers who have audio on — trending sound, a well-timed line, music that pays off — rather than leaving sound as decoration over a piece that already works muted.

    Spark Ads

    Spark Ads is TikTok's ad format for putting budget behind a post that already exists — yours or a creator's, with their authorization — so the paid version keeps the post's real likes, comments and shares instead of starting from zero.

    Stitch

    A Stitch is TikTok's clip-and-continue format: a few seconds of someone else's video, then yours. It is not the editing term for joining clips.

    Stitching a comment

    Stitching a comment is using a viewer's comment as the next hook: a reply-to-comment video, not TikTok's Stitch clip-and-continue.

    Synthetic media disclosure

    Synthetic media disclosure is the broader legal obligation to identify AI-generated or altered content as such — increasingly a binding requirement rather than a platform preference, with the EU AI Act's Article 50 as the clearest current example.

    Synthetic performer

    A synthetic performer, under N.Y. GBL §396-b (9 June 2026), is a generated human-looking performer who is not recognizable as an identifiable real person.

    Synthetic persona

    A synthetic persona is a fake human presenting as an authority — YouTube's July 2026 third inauthentic bucket — not merely a generated talking head.

    Testimonial ad

    A testimonial ad is a paid clip of someone endorsing a product — customer, creator, or avatar — and it still needs an endorsement disclosure.

    TrueView

    TrueView is YouTube's skippable in-stream ad: the viewer can skip after five seconds, and you pay on a qualified view, not every impression.

    UGC

    UGC meaning: user-generated content — photos, videos, or reviews from customers or creators, not the brand. In 2026 it also covers paid UGC-style ads.

    UGC brief

    A UGC brief is the shot list plus usage rights a creator or generator is hired to fill — deliverables and a licence, not a vibe about authenticity.

    UGC creator

    A UGC creator is paid to produce content in the UGC style for a brand's own use — with no requirement to post it to their own account — which makes the job closer to a content contractor than to an influencer.

    Unboxing

    An unboxing is a video of opening a product on camera — packaging, first look, first use — a UGC-native format that lives or dies on hands and the box.

    Usage rights

    Usage rights are the contract terms that say where a brand can run a creator's content, for how long, and in what form — organic-only, paid, whitelisted, in perpetuity — and they default to the narrowest reading if the contract doesn't say.

    Usage window

    A usage window is the timebox on paid use of a creator's content — how long the licence to run the file lasts, not whether they can work for a rival.

    Vibes remix

    A Vibes remix is a new AI video made from someone else's clip inside Meta's AI-only Vibes feed, then optionally cross-posted to Reels or Stories.

    Whitelist code

    A whitelist code is the advertiser token a creator issues so a brand can run that post as an ad from the creator's handle.

    Whitelisting

    Whitelisting is a brand running paid ads directly through a creator's own handle — with the creator's permission — so the ad shows the creator's name and profile instead of the brand's.

    Every term, A to Z

    4K native9:16 vs 4:5 vs 1:1A-rollAgent memoryAI content labelAI dubbingAI slopAI UGCAI-generated profileAlpha channelAny-to-anyArticle 50Aspect ratioASRAudio extendAudio tagsAudio-to-videoAutoregressive modelAverage view durationB-rollBackground removalBatch generationBeat listBilling typeBit depthBoosted postBrand kitBrand safetyBurned-in captionsCadenceCamera controlCamera pathCaption line lengthCaption timingCFG scaleCharacter consistencyCharacters per secondChroma subsamplingClosed captionsCodecCold openComment baitCompletion rateConspicuous disclosureContent calendarContent credentialsContent repurposingContent RewardsContent velocityCover songCPM vs. CPVCreative fatigueCreativity ProgramCreditCredit packCTA cardDark postingDeepfakeDemand genDenoising strengthDeployerDiegetic soundDiffusion modelDiffusion transformerDigital-twin creatorDistillationDriving clipDuetDurationElo ratingEnd screenEngaged viewsEvergreen vs trendExclusivity windowExpressive-work exemptionExtend retakeFaceless channelFine-tuningFirst-frame viewFirst-last frameFirst-second hookFlow matchingForced alignmentForced narrativeFrame interpolationFrame rateFTC endorsement disclosureGeneration timeGifting campaignHold rateHome feedHome Timeline impressionsHook packHook rateHook–body–CTAIdea PinImage-to-imageImage-to-videoInauthentic contentInpaintingJump cutKaraoke captionsLatent spaceLipsyncLoRALoudnessMaterial connectionMicro-influencerMotion controlMotion levelMulti-referenceMulti-scene generationMulti-speaker dialogueMultimodal modelMute autoplayMute-safeNative advertisingNative audioNative resolutionNegative promptNegative prompt (video)Off-screen speakerOne-tap templateOpen captionsOriginality scoreOutpaintingOverlay chromeOverlay labelPack shotPattern interruptPer-second pricingPhoneme-level lipsyncPhotoreal disclosurePhysics consistencyPicture-in-picturePlatform-native captionsPremade avatarProduct feedPromptPrompt adherencePrompt expansionPrompt weightingProvenanceProvider (AI Act)Qualified watch hoursReading speedRecurring seriesReference imageReference stackReference-to-videoRegion editRepost limitRepurposingResolutionRetention curveReusable characterRewatchRPMSafe areaSafe-titleSafety checkerSampling stepsSaved workflowSavesScene continuityScheduler stripSDHSeatSeedSeed lockSeedingSegmentationShare of voiceShares per viewSilent rowSlideshowSound-on strategySpark AdsSpeaker IDSpeech-to-speechSpeech-to-textSRTStem separationStill contractStitchStitching a commentStoryboardStyle presetSwipe-awaySyndicationSynthetic media disclosureSynthetic performerSynthetic personaSynthIDTalking headTemporal consistencyTestimonial adText overlayText-to-imageText-to-musicText-to-SFXText-to-speechText-to-videoThumbnailTrueViewTwo-line limitUGCUGC adUGC briefUGC creatorUnboxingUpscalingUsage rightsUsage windowVibes remixVideo extendVideo inpaintingVideo inputVideo-to-videoVisible watermarkVoice cloningVoice designVoice isolationVoice stabilityVoiceoverWatch timeWatermarkWebVTTWhitelist codeWhitelistingWord timestampWord-highlight captions

    Looking for something else?

    This is vocabulary, not instructions. For the settings a specific model exposes see the model catalog; for how to actually do a job see the editing tasks or the workflow recipes.