Two heads: generating past native resolution
A canvas far above native size duplicates subjects and limbs. Generate at the model's real ceiling, then upscale, and the test that finds that ceiling.
Ask an image model for a canvas far above the size it was trained around and it does not add detail. It tiles the composition it knows how to draw. You get a second head, a repeated torso, a third arm growing out of a waist, a window that copies itself down the wall. The file is large. The picture is two pictures overlapping.
This is not a seed problem and it is not a prompt-adherence problem. It is a spatial prior hitting a canvas it was never trained to fill in one pass. The reliable rule is boring and it holds: generate at the model's native size, pick the take, then upscale. The rest of this post is how to recognise that you are past the ceiling, and how to find that ceiling without burning a production batch on it.
What "too large" actually means
Resolution is how many pixels the output contains, and it is set before generation. Models also have a native size they were trained around. Those two numbers are not the same thing, and the spec sheet will not always tell you they differ.
A quality ladder that offers SD, HD and 4K is a menu of output sizes, not a promise that the 4K rung is generated natively. Two-pass regeneration versus native high resolution is the video version of this split. The image version has the same tell: repeating structure, extra subjects, and detail that looks sharp until you notice it is the same detail twice.
The canvas can go too large in two directions.
Too many megapixels. The model's spatial prior covers a certain field of view at a certain pixel density. Stretch both axes at once and it fills the extra room by repeating the subject it already solved. Two faces. Two hats. A body that starts again at the hips.
An extreme aspect ratio. A very tall 9:16 or a very wide 16:9 at an otherwise reasonable pixel count can tile along the long axis even when a 1:1 at the same megapixel count would not. Flux 1.1 Pro is a known case: very tall or very wide frames occasionally duplicate subject elements, and the practical fix is to generate closer to a standard ratio and outpaint. Same failure, different trigger. The prior ran out of unique composition along one axis and started again.
If you are seeing duplication, stop adding adjectives. "Single subject, one person, only one head" names the concept you are trying to suppress, and naming it does not allocate a larger spatial prior. Change the canvas.
How tiling shows up in the picture
Duplication is easy to miss at thumbnail size. A second head at the edge looks like a busy composition until you count.
Look for these, in this order:
- A second instance of the subject, often smaller, often along the long axis, often cropped by the edge.
- Repeated limbs on one body — a spare arm, a leg that belongs to no torso, a hand that is a copy of the other hand at a different scale.
- Tiled background structure — windows, columns, trees, shelves repeating at an interval the set would not have.
- A subject that starts again halfway down a tall frame: the same jacket, a second face, as if the portrait was drawn twice and stacked.
If two or more of those are present, you are past native, not getting a weird seed. Changing the seed at the same size will give you a different tiled picture. That is the diagnostic. If the error is compositional noise, seeds diversify it. If the error is tiling, seeds remix the tiles.
A crowded prompt can produce extra people for a different reason: you asked for a group and the model over-counted. Tiling is repetition of the same subject, not extra distinct people. A clone of your hero in the corner is tiling. A stranger in the corner is over-counting. Simplify the prompt for over-counting. Change the canvas for tiling.
Generate at native, then upscale
- Generate at the last size that does not tile. For most image models that is a standard ratio (1:1, 3:2, 2:3, 4:5) at the middle quality rung, not the top one. Iterate there. Composition, light, identity, wardrobe: all of that is cheaper to discover when the model is drawing one picture.
- Pick the take. Do not upscale a tiled image. The upscaler will faithfully enlarge the second head.
- Run a dedicated upscaler on the keeper. Asking the agent to upscale an image is that pass. You are asking a model whose whole job is adding texture, not one whose job is inventing a scene at a size it cannot hold.
- Check the upscale at 100 percent for invented information. Faces at a distance, small text, logos. Texture is allowed to be invented. Copy and identity are not. If a label changed wording, you needed a native generation at a tighter crop, not a bigger file.
This is the same decision upscale the render or generate at full resolution walks through for video. For stills the compressed version is: upscale when the detail that matters is texture; generate native and tighter when the detail that matters is a countable thing (one person, one product, one block of copy). Duplication is what happens when you ask the scene model to be the upscaler.
If you need a tall or wide delivery, do not start there. Generate at a standard ratio, then outpaint the extra canvas in one or two steps so the model extends the set instead of inventing a second subject. Outpainting a correct picture into a 9:16 is a different job from generating a 9:16 from noise, and it is the job that does not tile the hero.
Put the delivery size in the export, not in the prompt. "Medium shot, waist up" is a composition instruction. "8k full body" is a size instruction wearing a composition costume.
Finding a model's real ceiling
Do not guess, and do not trust the largest label on the model page as native. The catalog tells you the menu. A five-minute test tells you the ceiling.
Read the menu first. Open the model and note the quality rungs it actually exposes. Flux 1.1 Pro lists SD, HD and 4K; that is a ladder, not a native-size confession. Vocabulary is inconsistent across providers — 1K/2K/4K, HD/SD, exact pixel heights — and they do not map cleanly, which is why the resolution glossary tells you to read the model's own list.
Then run a ceiling strip. One subject, plain background, no crowd, no tiny text. Same prompt, same seed, each quality rung and each ratio you actually use (1:1, 4:5, 9:16, 16:9). You are looking for the first rung where tiling appears.
| Result | What it means | What you do |
|---|---|---|
| One subject, unique background, fine texture | Inside native, or close enough | Safe to iterate here |
| Soft but not repeating | Undersized, or a two-pass top rung that has not tiled yet | Usable; upscale if you need bite |
| Second subject, repeated limb, tiled set | Past native | Drop one rung, or change ratio, then upscale |
| Clean at 1:1, tiled at 9:16 | Ratio is the trigger, not megapixels | Generate 4:5 or 1:1, outpaint to 9:16 |
Pin the prompt and the seed while you change size. If you change wording and size together you will not know which one tiled the frame.
Five or six frames is enough. Once you have the last clean rung for that model, write it down next to the model name and stop rediscovering it on client work. The ceiling is per model. A model that holds 1:1 at its top rung can still tile a 9:16 at the same label.
If you dispatch through the agent, ask it for the model's input schema before you set a custom size. That is the allowed values for that endpoint, not a remembered number from a different model in the same family. View the keeper at 100 percent along each edge, then down the vertical centre of a tall frame. That is where the second head usually sits.
FAQ
Is a duplicated limb the same problem as a six-fingered hand?
No. A six-fingered hand is a small-region anatomy failure: the model averaged a pose space. A duplicated limb in an oversized canvas is the composition tiling. If the extra arm is a copy of the other arm, or a second torso is present, drop the size. If one hand has six fingers and the rest of the picture is a single clean subject, inpaint the hand and leave the canvas alone.
Can I inpaint the extra head and keep the large file?
You can, and you will do it again on the next large file. Inpainting a symptom of an oversized canvas is a tax on every take. Drop the size, regenerate, upscale the keeper. Save inpaint for defects that are actually local.
Does this happen in video too?
Yes, with a different costume. A video model asked for a wide, long, high-resolution shot will duplicate subjects or invent a second character at the edge, then animate the mistake. Generate at a comfortable size and ratio, then upscale the clip if the delivery needs it. Watch the upscale at full speed; per-frame invention on a tiled region reads as a person flickering into existence.
Why did a 1:1 at the top tier look fine and the 9:16 of the same prompt grow a second person?
The long axis ran past the prior. Generate 4:5 or 1:1, then outpaint. Or generate 9:16 at a lower rung. Extreme ratios and extreme megapixels are two ways to ask for a picture the model will fill by repeating itself.