A photo background cutout is one instruction, not a video key
generate_image_from_image isolates a still. Treating that cutout job as a clip key or a restyle wastes credits.
A photo background cutout is one instruction, not a video key. Attach a still, ask for transparent or a plain replacement, and generate_image_from_image routes an image-edit model to isolate the subject. That is the whole job. Using it on a clip, or stuffing a restyle into the same sentence, wastes credits on a file you will not ship.
Remove the background from a photo is the named capability. One instruction, a clean cutout back. The text-to-image tool is a different door: it starts from words, not from a photo you already have.
What this job actually is
You have a product still, a headshot, a pack shot. You want the subject on transparent, or on white, or on a backdrop you can name in a clause. You attach the photo. You say what the new back should be. The agent picks an image-edit model and returns a new still. remove_background can sit as a supporting call; the billed row is still an image edit, priced per model and shown before you confirm.
The output is a cutout you can drop into a slideshow, an ad, or a composite. It is not a keyed talking-head video. It is not a generated scene. It is not "make this look like a studio" if that sentence also changes wardrobe, lighting, and crop.
What wastes credits
Three mix-ups show up constantly.
You handed it a clip. Background removal on video is a different capability: remove the background from a video. That path keys solid-black talking-head plates in FFmpeg, or picks a video model for real scenes. Running the still job on a frame you hoped would become a keyed clip leaves you with one image and a video you still have to process.
You wanted a restyle. "Remove the background and make it cinematic with golden hour" is edit a photo with AI. Same primary tool family, different instruction. The cutout job is narrow: isolate, maybe replace the back. Extra adjectives spend the edit on a new picture you did not ask to keep.
You wanted a new product image from text. Do not attach a photo just to ignore it. Generate an image from text is that row.
A miss is a full edit charge for a file you throw away. Name the job: cutout, not "clean this up."
The test
Look at the file you are attaching.
If it is a still and the only ask is "subject out, back gone or replaced," this is the job. Stop adding style clauses.
If it is a video, leave this page. If the still is already fine and you want motion, that is image-to-video, not a transparent PNG. If you need the person on a new set with a new outfit, write a real photo-edit brief instead of a background-removal prompt with extra words.
Honesty about the file is the whole decision. This is one named agent job. Using it for a different job wastes credits.
FAQ
Is this the same as a dedicated background-removal model?
No. For stills the agent routes through Versely's image-edit models via generate_image_from_image. You still get a cutout. You do not get a separate SKU just because the sentence said "remove background."
Can I use this on a talking-head video?
Not as this job. Video is remove the background from a video. Extracting one frame, cutting it out, and hoping the rest of the clip follows is two wrong tools stacked.
Does a replacement backdrop still count as this job?
Yes, if the subject stays and the back is the only change. Once you also change clothes, lighting, or the pose, you have left the cutout job and entered a photo edit.
Why not just generate a new product shot from scratch?
You can, if you do not care about this product's actual pixels. The cutout exists because the still is already the contract. Regenerating from text is how you lose the pack, the label, and the lighting you already paid a photographer for.