Flux 2 Image Prompting: A Working Reference
Flux 2 stills prompting: sentence prompts, five-slot structure, photoreal cues. Run Pro, Max, Flex, and Flash on Versely text-to-image.
On Versely, Flux 2 prompts are sentences, not keyword piles. That single fact explains most of the gap between people who get magazine-grade output from the Flux 2 family and people who get generic renders: the model family was trained to follow natural language descriptions, so "cinematic, 8k, masterpiece, trending, ultra-detailed" (the incantation style inherited from older models) does close to nothing here, while a plain declarative sentence describing what's in the frame does nearly everything. This is a working reference for Flux 2 image prompting: the structure, the vocabulary that actually moves output, and the iteration loop that converges instead of wandering. Catalog rows on Versely: Flux 2 Pro at 3 credits, Max at 6 credits, Flex at 4 credits, Flash at 1 credit.
Write sentences, drop the incantations
Flux's prompt understanding rewards specificity delivered as prose. Compare the two philosophies on the same brief:
Keyword style: "woman, coffee shop, laptop, cozy, warm light, bokeh, 8k, photorealistic, masterpiece"
Sentence style: "A woman in her thirties with short curly hair works on a laptop at a window table in a small coffee shop. Morning light comes through the window behind her. Shot on a 50mm lens with shallow depth of field."
The second prompt gives Flux relationships — who is where, light comes from where, the camera does what — and relationships are what it renders well. Quality-bait words ("masterpiece," "8k," "award-winning") don't hurt much, but they occupy attention that concrete nouns and spatial prepositions would use better. If you're migrating from a keyword-native model, the biggest single upgrade is deleting the incantations and spending those words on one more concrete detail.
The five-slot structure
Nearly every strong Flux prompt fills five slots, usually in this order:
| Slot | What goes in it | Example |
|---|---|---|
| Subject | Who/what, with 2–3 concrete attributes | "an elderly potter with clay-stained hands" |
| Action/pose | What the subject is doing | "shaping a bowl on a spinning wheel" |
| Setting | Where, with one or two anchoring props | "in a cluttered sunlit studio, shelves of glazed pots behind" |
| Light | Source, direction, quality | "low afternoon sun through a side window, dust visible in the beam" |
| Camera/style | Lens, distance, medium | "medium close-up, 85mm, soft film-like color" |
You don't need every slot every time — but when a render disappoints, the diagnosis is almost always an empty slot the model filled with a default. Generic-looking light? You never specified light. Boring framing? Slot five was empty. The structure doubles as a debugging checklist.
Two refinements worth internalizing:
- Attributes beat adjectives. "Clay-stained hands" outperforms "realistic detailed hands" because it's a fact about the scene, not a quality plea.
- Prepositions are load-bearing. "Behind," "beside," "reflected in," "half in shadow" — Flux handles spatial relationships unusually well, and prompts that exploit that produce compositions other models can't hold together.
Photorealism cues that actually register
For commercial photorealistic work — the job Flux 1.1 Pro does daily — a small vocabulary earns most of the realism:
- Lens and distance: "35mm environmental portrait," "85mm close-up," "wide 24mm interior." Focal length language shifts perspective rendering, not just crop.
- Aperture effects: "shallow depth of field, background softly out of focus" — more reliable than the bare token "bokeh."
- Light quality: "overcast soft light," "hard noon sun with sharp shadows," "single warm practical lamp in a dark room." Naming the source makes shadows fall coherently.
- Imperfection: "slight film grain," "candid, mid-motion," "flyaway hairs." Perfect renders read as CGI; one imperfection cue lands the photo.
- Surface truth: "condensation on the glass," "worn leather grain," "fingerprints on the steel." Material details are cheap words with huge realism return.
What you can mostly skip: negative-style disclaimers ("no distortion, no extra fingers"). Flux handles anatomy well enough that positive description is the better spend — see the negative prompts guide for when exclusions do earn their place.
An iteration loop that converges
Random re-prompting is the main way people burn credits on Flux. This loop converges in three to five generations:
- Generate from the five-slot prompt. Judge the output against slots, not vibes: subject right? light right? framing right?
- Fix one slot per revision. Wrong mood → rewrite the light slot only. Wrong composition → camera slot only. Keep every other word identical so you can attribute the change.
- Lock and vary. Once four slots are right, re-roll a few takes — Flux's variance across seeds on a fixed prompt is your free art direction, and picking among takes beats forcing one take to be perfect.
- Upscale the winner. Do detail enhancement and 4K upscaling as a final pass on the selected image rather than prompting for "extreme detail" up front.
The discipline of one-change-per-revision feels slow and is actually fast: after ten prompts you know what each slot does in your niche, and your first drafts start landing. Teams that keep a shared bank of proven slot phrases get compounding returns — that practice is covered in prompt engineering for image generation.
Where Flux fits in your stack
Flux is the workhorse: prompt adherence, anatomy, and text handling good enough that it's the sensible default for product shots, lifestyle scenes, blog heroes, and ad stills. It's not the only tool — heavy typography favors a typography-first model, and extreme stylization sometimes wants a different family — but "start with Flux, leave for a reason" is a defensible policy. The economics reinforce it: Flux 1.1 Pro as a workhorse model breaks down why its cost-per-usable-image stays low for exactly the sentence-prompting reasons above. On Versely you can run the same five-slot prompt across Flux and two or three neighboring models side by side and let the outputs argue.
Skip the dialect: brief the Versely agent
You can learn the five-slot structure. On Versely you can also skip memorizing every Flux dialect and brief the agent instead of writing prompts: content type, photos or references, the edits you will accept, target platforms, and a budget ceiling. Name the row when it matters ("iterate Flux 2 Flash, final on Flux 2 Pro"). The agent plans the job; you approve the plan instead of babysitting slot order.
A prompt still wins when you are hand-tuning one hero still. A brief wins when the job is bigger than one generate: variants, cleanup, and a post path with a spend cap.
After the take: edit, post, collections
When the Flux 2 pass is close enough:
- Finish cutouts and sharpening on-device with /free-tools/background-remover and /free-tools/image-upscaler when those are the remaining gaps. Use /free-tools/object-eraser for stray props.
- Optional motion: send a winning still into /tools/ai-video-generator as image-to-video. Prompt the motion brief on the video row; do not ask the still model to invent a timeline.
- Upload or schedule to the social accounts Versely already connects for your workspace. If a network is not connected, export the PNG or JPEG and post from the native app.
- Save keepers into a Versely collection so the next brief reuses the same product refs, winning prompt lines, and approved stills.
Honest limit: Flux 2 is a stills family, not a timeline tool. Board separate generates for separate setups. Credits apply on every generate; Flash is for cheap iteration, Pro or Max for the keeper.
FAQ
How long should a Flux prompt be?
Two to five sentences covering subject, action, setting, light, and camera — typically 40–90 words. Shorter leaves slots for the model to fill with defaults; much longer usually means you're describing two images at once. Spend extra length on concrete details, never on quality adjectives.
Do keyword-style prompts work on Flux at all?
They produce output, but they waste the model's main strength: natural-language relationships between elements. Comma-separated keywords give Flux a bag of ingredients with no instructions. Full sentences with spatial prepositions consistently produce more coherent, art-directable images.
Why do my Flux images look generic?
Almost always an empty slot — usually light or camera — that the model filled with its statistical default. Specify a light source and direction plus a lens and distance, and add one imperfection or material detail. Generic output is a specification problem before it's a model problem.
Should I use negative prompts with Flux?
Rarely. Flux's anatomy and coherence are strong enough that positive description outperforms exclusion lists in most cases. Reserve exclusions for recurring, nameable intrusions — a watermark style, an unwanted object class — rather than defensive boilerplate on every prompt.
How do I keep a series of Flux images consistent?
Freeze your style, light, and camera slots verbatim and vary only subject and action between images. For character continuity across a series, add a fixed 2–3 attribute description block for the character and reuse it word-for-word in every prompt.
Take the five-slot template into Versely's text-to-image tool, run it on Flux 1.1 Pro, and fix one slot per revision — a $1 credit pack covers a full convergence loop.