

Alibaba’s Qwen team released Qwen-Image-2.1, a new open-source AI model focused on image generation and editing. The model combines text-to-image generation and image editing within a single architecture.
Its visual generation component contains 7 billion parameters and uses 32 Single-Stream Diffusion Transformer (DiT) layers. Alibaba says the architecture balances image quality with inference efficiency and computational cost.
A key addition in Qwen-Image-2.1 is native transparency. The model can generate RGBA images with transparent backgrounds directly from text prompts. Users can also edit transparent layers and extract subjects from existing photographs.
The new model also expands multi-image editing capabilities. Users can provide up to 10 reference images for a single editing or composition task. Allowing creators to combine several subjects, products or visual elements while preserving important characteristics.
Qwen-Image-2.1 supports native 2K resolution and offers multiple aspect ratios. Its recommended configurations include resolutions such as 2048×2048 for square images and 2752×1536 for 16:9 content. The model also focuses on improving typography, portrait lighting, textures and fine visual details.
Qwen also supports local editing through circles, painted annotations and separate masks. Thus, giving users greater control over specific areas of an image.
Alibaba introduced mixed-granularity attention and prefix KV-cache reuse to reduce repeated computation. These optimisations are for users who work with multiple reference images or perform repeated editing operations.
The model is available through platforms including Hugging Face and ModelScope. Qwen says it also received Day-0 support from tools such as Diffusers and ComfyUI. You also get inference frameworks including vLLM-Omni and SGLang.
Also Read: Alibaba Launches Qwen AI-Powered Smart Glasses to Challenge Meta’s Ray-Ban Model