GPT Image 2 Prompting Guide
Prompt GPT Image 2 as a stills brief: composition, quoted text, product lock, style. Run the ~9-credit row on Versely.

GPT Image 2 prompting is not a mood board caption. Treat the box like a stills brief: deliverable, subject, composition, verbatim text, style, and what must stay out. On Versely you run GPT Image 2 Text to Image at about 9 credits per generate (composer totals still move with resolution), then iterate the brief instead of restacking adjectives. There is no free generation tier on this row; credits apply.
This guide is the technique layer for stills. For row routing see when to pick GPT Image 2 text-to-image. For type-heavy jobs see GPT Image 2 versus Nano Banana 2 for legible type. Deeper vocabulary lives in shot composition prompts and style keywords. Video hubs (Seedance 2.5, Wan 3.0) cover motion; this post is the stills companion.
The stills brief stack
Only subject and deliverable are mandatory. Add the other slots in the order your last still failed.
- Deliverable - product photo, poster, thumbnail base, packaging comp, diagram
- Subject - who or what must stay readable
- Scene - place, materials, light direction, color
- Composition - framing, viewpoint, placement, aspect ratio, negative space
- Text - quoted strings, typography, line count, placement (or no text)
- Style - photorealistic trigger, medium, lens language, grade
- Constraints - exclusions and preserve lists (no dedicated negative field on this stack)
Write it as labeled lines if that helps you edit. GPT Image 2 reads long structured briefs like a chat model. What it does not forgive is five conflicting ads glued with commas.
Composition before adjectives
Name the shot size and viewpoint before you stack mood words. Call out placement when layout matters, and state the aspect ratio when the canvas is load-bearing (9:16, 4:5, 16:9, 1:1).
Weak:
A beautiful premium soda bottle looking luxurious and viral.
Stronger:
FORMAT: 4:5 product still for Instagram. SUBJECT: A frosted citrus soda bottle on wet slate. COMPOSITION: Bottle in the lower-right third. Empty upper-left for headline. Three-quarter view, eye level, shallow depth of field. LIGHT: Soft daylight from upper left, fresh commercial grade. STYLE: Photorealistic product photography, 50mm equivalent. CONSTRAINTS: No text, no logos, no watermark, no extra bottles.
Pair shot size with lens language (close-up, 85mm; wide, 24mm; macro for label texture). Two concrete composition words beat ten vague intensifiers.
Worked prompt 1 (hero product still)
DELIVERABLE: 1:1 ecommerce hero, white seamless.
SUBJECT: Matte navy ceramic mug with a cream handle, three-quarter view.
COMPOSITION: Centered. Soft contact shadow only. Room for a price badge in the lower-right corner (leave that corner empty).
LIGHT: Softbox from upper left, gentle fill from the right.
STYLE: Photorealistic studio product photography, sharp focus on the lip and handle join.
CONSTRAINTS: No text, no logo, no watermark, no extra props, no people.
Text in the image: quote it, then fence it
GPT Image 2 is strong at readable in-image type when you treat copy as literal output. Three habits:
- Put required wording in straight quotes.
- Specify typography and placement in a separate clause (font feel, size, color, zone).
- Say
render verbatim, exactly as written, no extra charactersandno other text anywhere in the frame.
For multi-line layouts, count the lines. Spell unusual brand names letter-by-letter the first time. Dense menus and packaging panels want higher fidelity when the composer exposes quality or resolution controls; iterate cheaper only while you lock layout.
Weak: A poster that says something about summer vibes.
Stronger:
DELIVERABLE: 9:16 event poster.
SCENE: Single silhouetted figure on an empty street under one lamp at twilight. Deep blues and warm amber.
TEXT: Large serif gold headline across the lower third reads "THE NIGHT BEGINS AT EIGHT". Smaller sans-serif white tagline below reads "A STORY ABOUT WAITING". Two lines only. Render every line verbatim, exactly as written, no extra characters.
CONSTRAINTS: No other text, no watermark, no logos, no QR codes.
If type is the whole job, also read the legible type comparison. If captions belong on a video later, burn them after with /tools/ai-caption-generator or /free-tools/burn-captions instead of asking the still to invent UI chrome.
Product stills and reference locks
When you attach a product photo, uploading is not the instruction. Label the asset (Image 1) and give it one job. State a preserve list for SKU geometry: shape, cap, label art, proportions, color. State what not to copy from the reference (studio backdrop, lighting, props).
REFERENCE: Image 1 is the perfume bottle to place. Preserve the bottle's shape, cap, label, and exact proportions. Do not alter the bottle. Do not copy Image 1's white seamless background or lighting.
SCENE: Moody marble bathroom counter, folded white linen towel, one eucalyptus sprig. Soft window light from the left, slight steam in the background.
COMPOSITION: Bottle sits left of center on the counter. Three-quarter view.
STYLE: Photorealistic editorial product photography.
CONSTRAINTS: No extra bottles, no rewritten label text, no watermark.
Edit prompts work the same way: Change only X plus an explicit preserve list. Iterate one change at a time and restate the preserve list on every follow-up.
Style control that actually sticks
Say photorealistic when you want realism; it is the strongest single trigger on this family. Add camera language for composition (lens, depth of field, light direction), not for fake EXIF precision. Keep the style block to three to six concrete terms after the subject. Prefer lighting, lens, medium, and era words over empty intensifiers.
Test one keyword at a time on a boring base subject (style keywords). Subject first, style after.
Worked prompt 2 (lifestyle still)
DELIVERABLE: 4:5 lifestyle still for a coffee brand.
SUBJECT: A barista in a navy apron pours from a gooseneck kettle into a paper filter.
SCENE: Narrow cafe counter at sunrise. Warm window light from the left. Matte ceramic and brushed steel.
COMPOSITION: Medium shot on hands and kettle, 50mm feel, shallow depth of field.
STYLE: Photorealistic documentary still, soft contrast, warm earth grade.
CONSTRAINTS: No text, no logo, no watermark, no extra people, no warped hands.
When to upscale or remove background after
Do not burn another generate to fix a soft edge or a busy backdrop if a post step will do.
- Remove background: approved hero needs a clean PNG. On-device /free-tools/background-remover (browser; no upload).
- Upscale: composition and type are right, still is soft for a large crop. On-device /free-tools/image-upscaler (WebGPU Anime4K when available; honest canvas fallback otherwise).
- Object cleanup: stray props via /free-tools/object-eraser on-device.
Regenerate when the brief is wrong (composition, copy, product geometry, style). Free tools clean pixels; they do not rewrite a contradictory brief.
Failure modes
| Failure | Likely cause | Fix in the brief |
|---|---|---|
| Centered generic catalog look | No composition or aspect | Name framing, placement, ratio, negative space |
| Rewritten or gibberish type | Unquoted copy, no verbatim fence | Quote strings; typography clause; no other text |
| Product redesigned mid-edit | Vague "use my image" | Preserve shape, cap, label, proportions; say what not to copy |
| Style washes out | Six conflicting mood words | Three to six concrete style terms after subject |
| Soft still or busy backdrop after approval | Asking the model to "fix" cleanup | On-device /free-tools/image-upscaler or /free-tools/background-remover |
| Watermark or fake UI | Web-trained chrome leak | Constraints: no text, no watermark, no logos |
| Empty upper zone filled anyway | Placement not protected | Reserve the zone: no subject in upper third |
Exclusions live in the prompt text. Targeted absences (no text, no people, no watermark) beat long boilerplate stacks; see also negative prompts.
Skip the dialect: brief the Versely agent
You can learn this stack. On Versely you can also skip memorizing every GPT Image dialect and brief the agent instead of writing prompts: content type, photos or references, edits you will accept, target platforms, and a budget ceiling. Name the row when it matters ("use GPT Image 2 Text to Image"). The agent plans the job; you approve the plan.
A prompt still wins when you are hand-tuning one hero still. A brief wins when the job is bigger than one generate: variants, cleanup, and a post path with a spend cap.
After the take: edit stills, post, collections
When the GPT Image 2 pass is close enough:
- Finish cutouts and sharpening on-device with /free-tools/background-remover and /free-tools/image-upscaler when those are the remaining gaps.
- Optional motion: send a winning still into /tools/ai-video-generator as image-to-video. Prompt the motion brief on the video row; do not ask the still model to invent a timeline.
- Upload or schedule to the social accounts Versely already connects for your workspace. If a network is not connected, export the PNG or JPEG and post from the native app.
- Save keepers into a Versely collection so the next brief reuses the same product refs, winning prompt lines, and approved stills.
Honest limit: GPT Image 2 is a stills generator, not a timeline tool and not a free unlimited sandbox. Board separate generates for separate setups. Credits apply on every generate.
FAQ
What is the best GPT Image 2 prompt formula?
Deliverable and subject first, then scene, composition, quoted text, style, and constraints (plus a preserve list when a reference is attached). Drop any slot you do not need. Add slots in the order your last failure showed up.
How do I get readable text in the image?
Quote the exact string, specify typography and placement, count the lines, and fence with verbatim / no extra characters / no other text. Check spelling in the output before you ship.
How should I use a product reference?
Address it as Image 1, assign one job, and list what to preserve and what not to copy. Product composites fail when the model is free to "improve" the SKU.
Is GPT Image 2 free on Versely?
No. The GPT Image 2 Text to Image row meters in credits (about 9 on the common stills path; totals still move with resolution). There is no free generation tier. On-device /free-tools handle remove-bg and upscale after you already have a still.
Where do I run this on Versely?
Open the text-to-image tool with GPT Image 2 selected, or start from the model page. Animate approved winners later in the AI video generator via image-to-video.
How is this different from Seedance or Wan prompting?
Those are motion briefs. This guide is stills: composition, type, product lock, post-cleanup. Generate the still here; animate on a video row only after you approve the frame.
Takeaway
Prompt GPT Image 2 like a short stills brief, not like a vibe caption. Lock composition and aspect, quote every load-bearing string, preserve product geometry when a reference is attached, and keep style concrete. Run the ~9-credit row on Versely, or hand the agent a job brief when you would rather not memorize the dialect. Clean winners on-device when upscale or remove-bg is enough, animate keepers via image-to-video when motion is next, and file keepers in a collection so the next still starts warmer.