Qwen-Image-2.1 generates a real alpha channel
Qwen-Image-2.1, released 20 September 2026, generates native RGBA with a true alpha channel up to 2K from a prompt, replacing generate-then-background-remove.
Qwen-Image-2.1, from Alibaba, generates native RGBA with a true alpha channel, up to 2K, straight from a prompt. It was released on 20 September 2026. The clear pixels are in that result. That replaces the generate-then-background-remove step.
The release account this page follows is what Alibaba actually shipped on 20 September. This page is the native alpha channel: RGBA up to 2K, cutouts, and edits that keep a clear background. It is not the Qwen Research License, not local VRAM and render time, and not composition from up to ten reference images.
Native RGBA is the output
Native RGBA means the checkpoint writes red, green, blue, and a true alpha channel in one result, up to 2K, straight from a prompt. The general term on this site is an alpha channel. On this model the channel is generated with the picture.
Generate-then-background-remove is the two-pass workflow this output replaces. A first pass generates a full frame. A second pass removes the background. Background removal is that second pass. The prompt path on Qwen-Image-2.1 returns the subject with the clear pixels already stored.
The visual generator is 7B, a 32-layer single-stream diffusion transformer. The text encoder is Qwen3-VL 8B. The VAE is a 64-channel RGBA VAE with 16x spatial compression.
| Piece | Specification |
|---|---|
| Visual generator | 7B, 32-layer single-stream diffusion transformer |
| Text encoder | Qwen3-VL 8B |
| VAE | 64-channel RGBA, 16x spatial compression |
| Prompt output | Native RGBA, true alpha channel, up to 2K |
Weights for the checkpoint went up the same day on Hugging Face, ModelScope and GitHub. Day-zero support is in ComfyUI, Diffusers, vLLM-Omni and SGLang.
One checkpoint, including Qwen-Image-Layered
Qwen-Image-2.1 unifies text-to-image and image editing in one checkpoint. It folds in Qwen-Image-Layered, which used to be a separate model.
A generated subject with clear pixels, and an edit that keeps a clear background, both come from that checkpoint. Text-to-image and image editing share the weights, with Qwen-Image-Layered folded into them.
Four jobs on a transparent layer
Apidog on what Qwen-Image-2.1 is, DataNorth on the release and eesel.ai on Qwen-Image-2.1 cover the same release. The transparency jobs in that coverage are four.
Generate from a prompt. Native RGBA, a true alpha channel, up to 2K. This is the job that replaces generate-then-background-remove. The 2K ceiling is stated for that generate. The other three jobs are listed without a separate resolution figure.
Edit a transparent image. The file is already transparent. The edit preserves the clear background.
Pull a subject from an ordinary photo. The input is an ordinary photo. The output is that subject as a transparent cutout.
Edit text inside a transparent layer. The layer is transparent. The text inside it is what the edit changes.
| Job | Starting point | Alpha result |
|---|---|---|
| Generate | A prompt | Native RGBA, true alpha, up to 2K |
| Edit | A transparent image | Clear background preserved |
| Cutout | An ordinary photo | Subject as a transparent cutout |
| Text | A transparent layer | Text edited inside that layer |
The clear pixels are a channel in the file the checkpoint writes.
A cutout keeps the colour cast
A true alpha channel marks which pixels are clear. The opaque pixels still carry the colour the model wrote.
Coverage of the release, including the eesel.ai review, records a yellowish colour tone and graininess, read as heavy GPT-Image training-data influence.
That cast is a composite problem. Clear pixels drop out when the cutout is placed on a new background. The opaque pixels stay. A prompt generate that came out as RGBA can still be yellow and grainy on the subject. A subject pulled from an ordinary photo can carry the same tone onto the cutout. Replacing generate-then-background-remove removes a pass. The alpha marks the clear pixels. The yellowish tone and the graininess travel with the pixels that remain.
Text in the layer is contested
Editing text inside a transparent layer is a listed job. Clean small text is a separate question, and the coverage leaves it contested.
The model struggles with infographics. It adds unsolicited Chinese text unprompted. Text rendering is contested: some call it much better, others still rate Ideogram superior for clean small text. The text encoder gets overloaded with larger prompts on complex compositional requests.
Those faults stay in the pixels a composite keeps. The struggle with infographics remains when the surround is transparent. Unsolicited Chinese text on a cutout composites with the subject. Others still rate Ideogram superior for clean small text when those letters sit inside a transparent layer. A larger prompt still overloads the text encoder on a complex compositional request.
Testers scored 7/15 versus 4/15
Community testers scored Qwen-Image-2.1 7/15 versus 4/15 for the prior model. The improvement is real. The same testing shows visible synthetic-data artefacts. The yellowish colour tone and the graininess sit with those artefacts, on a full frame and on the opaque pixels of a cutout.
No reproducible numerical benchmark table was published against closed rivals in the initial coverage wave. The 7/15 versus 4/15 figures are a community score against the prior model. Native RGBA up to 2K is a capability the 20 September release describes. A published numerical table against closed image models was not part of that initial coverage.
That is the 20 September checkpoint: native RGBA up to 2K, Qwen-Image-Layered folded in, and a community score of 7/15 versus 4/15 with the yellowish cast still on the cutout.