A mask is not an edit. It is a stencil: a black-and-white record of where a thing is, frame by frame. Everything interesting happens afterwards, when something else uses that stencil to delete, replace, blur, recolour or track the region it covers.
You point at what you want rather than describe it. Implementations take a click, a box drawn around the subject, or occasionally a text label, and propagate that selection forward through the clip. Propagation is the hard part — the object turns, is occluded, leaves frame and comes back.
It underpins several things that look like separate features. Background removal is segmentation of the subject plus a delete. Object replacement is segmentation plus inpainting. Tracking a logo across a shot is segmentation with the mask handed to a compositor.
In practice
- Selections are propagated, not re-detected per frame, so an early bad selection poisons the whole clip.
- Occlusion is the standard failure: a hand passing over the subject often splits the mask.
- Masks compose — several objects can be segmented separately and combined.
The mistake to avoid
Expecting a mask to be an output you can publish. It is an intermediate; on its own it is a silhouette.
Where you will run into it
- Remove a Video's Background — Two tools, two very different backgrounds.
Related terms
Background removal
Background removal separates a subject from everything behind it and discards the rest, leaving a cutout you can place on any backdrop.
Video-to-video
Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.
Inpainting
Inpainting regenerates a region you have masked while leaving the rest of the picture untouched, so a change stays local.
Temporal consistency
Temporal consistency is how well a generated clip keeps things the same from one frame to the next — a shirt that stays the same colour, a background that stays put.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.