What it produces is a matte — a per-pixel record of how much of each pixel belongs to the subject. Edges are rarely all-or-nothing, so a good matte is grey at the boundary, which is what lets hair and motion blur composite without a hard fringe.
On video the job runs per frame and has to agree with itself across the sequence. A matte that is a pixel wider on frame twelve than on frame eleven reads as a crawling edge, and that is the artefact that gives away a machine cutout more than any single frame ever does.
It replaces a green screen rather than competing with it. A physical green screen still gives a cleaner key on fine detail, but it needs the shot to have been planned that way; removal works on footage that already exists.
In practice
- Hair, glass, smoke and motion blur are the hard cases — everything else is close to solved.
- Subject-background contrast matters more than resolution: a dark coat on a dark wall is the worst case.
- Keep the cutout in a format that carries transparency, or the removal is undone on export.
The mistake to avoid
Checking the cutout against a white page. Composite it over the actual destination background — fringing that is invisible on white is obvious on dark.
Where you will run into it
- Remove a Video's Background — Two tools, two very different backgrounds.
- Remove the Background From a Photo — Same job, different pipeline than video.
- Add Picture-in-Picture to a Video — Reaction footage, corner-mounted on your main clip.
Related terms
Segmentation
Segmentation labels which pixels in each frame belong to a chosen object, producing a mask that other tools then act on.
Image-to-image
Image-to-image takes a picture as its primary input and returns a changed picture — restyled, corrected or varied — instead of inventing one from nothing.
Inpainting
Inpainting regenerates a region you have masked while leaving the rest of the picture untouched, so a change stays local.
Video-to-video
Video-to-video takes finished footage in and returns altered footage, using the original clip as the structural reference for every frame.
Text-to-video
Text-to-video is generation from a written prompt alone — you describe a shot, the model invents every frame of it, and no image or footage goes in.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.