Guides

    GPT Image 2 Prompting Guide

    Prompt GPT Image 2 as a stills brief: composition, quoted text, product lock, style. Run the ~9-credit row on Versely.

    Versely Team10 min read

    Versely Imagine GPT Image model picker

    GPT Image 2 prompting is not a mood board caption. Treat the box like a stills brief: deliverable, subject, composition, verbatim text, style, and what must stay out. On Versely you run GPT Image 2 Text to Image at about 9 credits per generate (composer totals still move with resolution), then iterate the brief instead of restacking adjectives. There is no free generation tier on this row; credits apply.

    This guide is the technique layer for stills. For row routing see when to pick GPT Image 2 text-to-image. For type-heavy jobs see GPT Image 2 versus Nano Banana 2 for legible type. Deeper vocabulary lives in shot composition prompts and style keywords. Video hubs (Seedance 2.5, Wan 3.0) cover motion; this post is the stills companion.

    The stills brief stack

    Only subject and deliverable are mandatory. Add the other slots in the order your last still failed.

    1. Deliverable - product photo, poster, thumbnail base, packaging comp, diagram
    2. Subject - who or what must stay readable
    3. Scene - place, materials, light direction, color
    4. Composition - framing, viewpoint, placement, aspect ratio, negative space
    5. Text - quoted strings, typography, line count, placement (or no text)
    6. Style - photorealistic trigger, medium, lens language, grade
    7. Constraints - exclusions and preserve lists (no dedicated negative field on this stack)

    Write it as labeled lines if that helps you edit. GPT Image 2 reads long structured briefs like a chat model. What it does not forgive is five conflicting ads glued with commas.

    Composition before adjectives

    Name the shot size and viewpoint before you stack mood words. Call out placement when layout matters, and state the aspect ratio when the canvas is load-bearing (9:16, 4:5, 16:9, 1:1).

    Weak:

    A beautiful premium soda bottle looking luxurious and viral.

    Stronger:

    FORMAT: 4:5 product still for Instagram. SUBJECT: A frosted citrus soda bottle on wet slate. COMPOSITION: Bottle in the lower-right third. Empty upper-left for headline. Three-quarter view, eye level, shallow depth of field. LIGHT: Soft daylight from upper left, fresh commercial grade. STYLE: Photorealistic product photography, 50mm equivalent. CONSTRAINTS: No text, no logos, no watermark, no extra bottles.

    Pair shot size with lens language (close-up, 85mm; wide, 24mm; macro for label texture). Two concrete composition words beat ten vague intensifiers.

    Worked prompt 1 (hero product still)

    DELIVERABLE: 1:1 ecommerce hero, white seamless.
    SUBJECT: Matte navy ceramic mug with a cream handle, three-quarter view.
    COMPOSITION: Centered. Soft contact shadow only. Room for a price badge in the lower-right corner (leave that corner empty).
    LIGHT: Softbox from upper left, gentle fill from the right.
    STYLE: Photorealistic studio product photography, sharp focus on the lip and handle join.
    CONSTRAINTS: No text, no logo, no watermark, no extra props, no people.
    

    Text in the image: quote it, then fence it

    GPT Image 2 is strong at readable in-image type when you treat copy as literal output. Three habits:

    1. Put required wording in straight quotes.
    2. Specify typography and placement in a separate clause (font feel, size, color, zone).
    3. Say render verbatim, exactly as written, no extra characters and no other text anywhere in the frame.

    For multi-line layouts, count the lines. Spell unusual brand names letter-by-letter the first time. Dense menus and packaging panels want higher fidelity when the composer exposes quality or resolution controls; iterate cheaper only while you lock layout.

    Weak: A poster that says something about summer vibes.

    Stronger:

    DELIVERABLE: 9:16 event poster.
    SCENE: Single silhouetted figure on an empty street under one lamp at twilight. Deep blues and warm amber.
    TEXT: Large serif gold headline across the lower third reads "THE NIGHT BEGINS AT EIGHT". Smaller sans-serif white tagline below reads "A STORY ABOUT WAITING". Two lines only. Render every line verbatim, exactly as written, no extra characters.
    CONSTRAINTS: No other text, no watermark, no logos, no QR codes.
    

    If type is the whole job, also read the legible type comparison. If captions belong on a video later, burn them after with /tools/ai-caption-generator or /free-tools/burn-captions instead of asking the still to invent UI chrome.

    Product stills and reference locks

    When you attach a product photo, uploading is not the instruction. Label the asset (Image 1) and give it one job. State a preserve list for SKU geometry: shape, cap, label art, proportions, color. State what not to copy from the reference (studio backdrop, lighting, props).

    REFERENCE: Image 1 is the perfume bottle to place. Preserve the bottle's shape, cap, label, and exact proportions. Do not alter the bottle. Do not copy Image 1's white seamless background or lighting.
    SCENE: Moody marble bathroom counter, folded white linen towel, one eucalyptus sprig. Soft window light from the left, slight steam in the background.
    COMPOSITION: Bottle sits left of center on the counter. Three-quarter view.
    STYLE: Photorealistic editorial product photography.
    CONSTRAINTS: No extra bottles, no rewritten label text, no watermark.
    

    Edit prompts work the same way: Change only X plus an explicit preserve list. Iterate one change at a time and restate the preserve list on every follow-up.

    Style control that actually sticks

    Say photorealistic when you want realism; it is the strongest single trigger on this family. Add camera language for composition (lens, depth of field, light direction), not for fake EXIF precision. Keep the style block to three to six concrete terms after the subject. Prefer lighting, lens, medium, and era words over empty intensifiers.

    Test one keyword at a time on a boring base subject (style keywords). Subject first, style after.

    Worked prompt 2 (lifestyle still)

    DELIVERABLE: 4:5 lifestyle still for a coffee brand.
    SUBJECT: A barista in a navy apron pours from a gooseneck kettle into a paper filter.
    SCENE: Narrow cafe counter at sunrise. Warm window light from the left. Matte ceramic and brushed steel.
    COMPOSITION: Medium shot on hands and kettle, 50mm feel, shallow depth of field.
    STYLE: Photorealistic documentary still, soft contrast, warm earth grade.
    CONSTRAINTS: No text, no logo, no watermark, no extra people, no warped hands.
    

    When to upscale or remove background after

    Do not burn another generate to fix a soft edge or a busy backdrop if a post step will do.

    Regenerate when the brief is wrong (composition, copy, product geometry, style). Free tools clean pixels; they do not rewrite a contradictory brief.

    Failure modes

    Failure Likely cause Fix in the brief
    Centered generic catalog look No composition or aspect Name framing, placement, ratio, negative space
    Rewritten or gibberish type Unquoted copy, no verbatim fence Quote strings; typography clause; no other text
    Product redesigned mid-edit Vague "use my image" Preserve shape, cap, label, proportions; say what not to copy
    Style washes out Six conflicting mood words Three to six concrete style terms after subject
    Soft still or busy backdrop after approval Asking the model to "fix" cleanup On-device /free-tools/image-upscaler or /free-tools/background-remover
    Watermark or fake UI Web-trained chrome leak Constraints: no text, no watermark, no logos
    Empty upper zone filled anyway Placement not protected Reserve the zone: no subject in upper third

    Exclusions live in the prompt text. Targeted absences (no text, no people, no watermark) beat long boilerplate stacks; see also negative prompts.

    Skip the dialect: brief the Versely agent

    You can learn this stack. On Versely you can also skip memorizing every GPT Image dialect and brief the agent instead of writing prompts: content type, photos or references, edits you will accept, target platforms, and a budget ceiling. Name the row when it matters ("use GPT Image 2 Text to Image"). The agent plans the job; you approve the plan.

    A prompt still wins when you are hand-tuning one hero still. A brief wins when the job is bigger than one generate: variants, cleanup, and a post path with a spend cap.

    After the take: edit stills, post, collections

    When the GPT Image 2 pass is close enough:

    1. Finish cutouts and sharpening on-device with /free-tools/background-remover and /free-tools/image-upscaler when those are the remaining gaps.
    2. Optional motion: send a winning still into /tools/ai-video-generator as image-to-video. Prompt the motion brief on the video row; do not ask the still model to invent a timeline.
    3. Upload or schedule to the social accounts Versely already connects for your workspace. If a network is not connected, export the PNG or JPEG and post from the native app.
    4. Save keepers into a Versely collection so the next brief reuses the same product refs, winning prompt lines, and approved stills.

    Honest limit: GPT Image 2 is a stills generator, not a timeline tool and not a free unlimited sandbox. Board separate generates for separate setups. Credits apply on every generate.

    FAQ

    What is the best GPT Image 2 prompt formula?

    Deliverable and subject first, then scene, composition, quoted text, style, and constraints (plus a preserve list when a reference is attached). Drop any slot you do not need. Add slots in the order your last failure showed up.

    How do I get readable text in the image?

    Quote the exact string, specify typography and placement, count the lines, and fence with verbatim / no extra characters / no other text. Check spelling in the output before you ship.

    How should I use a product reference?

    Address it as Image 1, assign one job, and list what to preserve and what not to copy. Product composites fail when the model is free to "improve" the SKU.

    Is GPT Image 2 free on Versely?

    No. The GPT Image 2 Text to Image row meters in credits (about 9 on the common stills path; totals still move with resolution). There is no free generation tier. On-device /free-tools handle remove-bg and upscale after you already have a still.

    Where do I run this on Versely?

    Open the text-to-image tool with GPT Image 2 selected, or start from the model page. Animate approved winners later in the AI video generator via image-to-video.

    How is this different from Seedance or Wan prompting?

    Those are motion briefs. This guide is stills: composition, type, product lock, post-cleanup. Generate the still here; animate on a video row only after you approve the frame.

    Takeaway

    Prompt GPT Image 2 like a short stills brief, not like a vibe caption. Lock composition and aspect, quote every load-bearing string, preserve product geometry when a reference is attached, and keep style concrete. Run the ~9-credit row on Versely, or hand the agent a job brief when you would rather not memorize the dialect. Clean winners on-device when upscale or remove-bg is enough, animate keepers via image-to-video when motion is next, and file keepers in a collection so the next still starts warmer.