Guides

    generate_images makes a still, not a clip

    generate_images is one still job. Routing a photo edit or a video ask through it burns the wrong credits.

    Versely Team4 min read

    generate_images makes a still. It is not a photo edit, not a video, not a talking head. Ask it for motion and you pay an image model to refuse a job it was never given.

    Generate an image from text is the named agent job. Describe a picture in English. The agent asks about style, aspect ratio, and count before it spends anything — it is instructed not to guess those and burn credits on the wrong look. Then it picks an image model (or uses the one you name) and returns stills to generation history. That is the whole contract.

    The text-to-image tool is the same door without the chat. Either way, the unit you bought is a picture.

    What the job is

    A prompt in. One or more images out. Required inputs on the tool are a prompt and a model. Optional extras — seed, resolution, negative prompt, reference images — still leave you in still-land. If you attach a photo because you want the photo changed, you have left this job. Image-edit models that require a starting frame live on a different row: edit a photo with AI. Passing a reference into generate_images because the catalog has a few edit-capable names does not make this the edit job. It makes you one mis-routed call away from a model that needed image_urls and did not get them, or a model that ignored the attachment and invented a cousin of your subject.

    Count is a stills decision. Four variations of a logo is four images. It is not a storyboard that will later "just become video." Lock the still first. Motion is a second bill.

    What it is not

    Three asks get stuffed into this row because they sound like "make me a picture of…":

    • A clip. Generate a video from text is generate_videos. Duration, aspect, resolution, audio. An image credit will not interpolate frames because you wrote "cinematic drone shot" in an image prompt.
    • An animated still. You already have the photo. Turn a photo into a video is generate_video_from_image. Starting from a blank prompt and hoping the image model invents a frame you can later animate is two jobs pretending to be one.
    • A cutout or a restyle of a file you are holding. That is the edit path. Background removal on a photo is its own capability. Do not spend a generate-from-text pass to "draw it again without the wall."

    Each of those has a priced model behind it. Using generate_images as a stand-in does not save the later bill. It adds a discarded still.

    The test

    If there is no source file, and the output you will ship is a picture — poster, product plate, thumbnail, reference for a later video — this is the job.

    If there is a source file you need respected, or the file you will upload is an MP4, stop. Name the other capability. Confirm the model. Spend once.

    Describe it. Let the agent pick an image model. Download the still. That is generate an image from text. Everything else is a different named job.

    FAQ

    Can I prompt an image model as if it were a video storyboard?

    You can write the sentences. You will get stills. A storyboard you actually shoot on a video model is a video job, priced per clip. Do not treat four images as a free animatic.

    Why does the agent ask about style and aspect before generating?

    Because generate_images charges per model, shown before you confirm. A one-line prompt with the wrong ratio is a paid miss. The follow-up is the cheap step.

    Is attaching a reference image the same as editing a photo?

    No. Edit-capable models on this tool still require the attachment to be passed correctly, and the named edit job exists so you do not have to remember which catalog rows demand image_urls. If the brief is "change this photo," use the edit capability.

    Does this replace the text-to-image studio?

    It is the agent form of the same still job. The studio is a launcher. The capability is the named task the agent is allowed to run. Neither one becomes video because the prompt mentioned a camera move.