About Qwen Image 2.1
Qwen Image 2.1 is the open-weight image model from Alibaba's Qwen team that folds text-to-image generation and image editing into one 7B-parameter diffusion transformer, paired with a Qwen3-VL 8B text and vision encoder.
It generates natively at 2K, 2048x2048 by default, across seven aspect ratios up to 2752x1536, and is the first Qwen image model with native transparent RGBA output, so it can produce cutouts and layered assets directly. It keeps the family's strong text rendering and identity preservation for people and products. With no input images it generates from the prompt; with up to ten reference images it edits or combines them, preserving identity and supporting local edits. Fine-tuning trains a LoRA on the text-to-image task from your own image and caption pairs.
Weights are released under the Qwen Research License, so commercial self-hosting needs a separate license from Qwen.
| Metric | Value |
|---|---|
| Parameter Count | 7B (plus 8B encoder) |
| Mixture of Experts | No |
| Max Resolution | 2752x1536 (2K) |
| Multilingual | Yes (Chinese, English) |
| Quantized* | No |
*Quantization is specific to the inference provider and the model may be offered with different quantization levels by other providers.
Ready to build with Qwen Image 2.1?
Try Qwen Image 2.1 in the Workbench to prompt it, compare outputs, and iterate on prompts without writing any code. When you're ready to ship, call the same model from our API and build your own apps on top of it.
curl -sSf -X POST https://hub.oxen.ai/api/ai/images/edit \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OXEN_API_KEY" \
-d '{
"model": "qwen-image-2-1",
"prompt": "A retro travel poster reading \"OXEN\" in bold condensed type, an ox on a mountain ridge at golden hour, clean vector illustration with a balanced layout",
"aspect_ratio": "16:9",
"resolution": "2K",
"guidance_scale": 1,
"output_format": "png"
}'Request parameters
Every field Qwen Image 2.1 accepts in the request body. See the API reference for response formats.
https://hub.oxen.ai/api/ai/images/edit| Field | Type | Required | Default | Description |
|---|---|---|---|---|
prompt | string | Yes | — | Text prompt describing the image to generate, or the edit to apply. Use @Image1, @Image2, etc. to reference the input images. Ask for a transparent background to get an RGBA cutout. |
input_images | array<string>nullable | No | — | Optional reference images to edit or combine, in order. Reference them in the prompt as @Image1, @Image2, etc. Up to 10 images. |
aspect_ratio | string | No | "16:9" | 1:116:99:164:33:43:22:3 |
resolution | string | No | "2K" | 1K2K |
guidance_scale | number | No | 1 | Classifier-free guidance strength, 1 to 10. 1 is the provider default with no guidance; raise it for sharper, more literal results. The negative prompt only applies above 1. Range: 1 – 10 |
negative_prompt | string | No | — | What to keep out of the image. Only applies when guidance scale is above 1. |
image_size | object | No | — | Pixel dimensions sent to fal, derived from aspect_ratio and resolution. |
output_format | string | No | "png" | Output format. PNG and WebP preserve transparency; JPEG fills it with white.pngjpegwebp |