Qwen Image — 6 NFE

A six-step distilled Qwen-Image-2.1 Generator for text-to-image and image editing. This repository contains the complete Transformer weights (297 tensors, FP32, 28.46 GB) in eight standard Diffusers safetensors shards, not a LoRA or a training-state checkpoint. The included runner automatically obtains the text encoder, VAE, processor and scheduler from the pinned official base model.

Quick start

Use Python 3.10+ and a CUDA-compatible PyTorch installation. The runner uses BF16. If this repository is private, first run hf auth login using an account with access. Do not put access tokens in scripts or notebooks.

pip install -U huggingface_hub
hf download Sutoonq/Qwen-Image-2.1-6NFE-Distill inference.py requirements.txt config.json model_index.json --local-dir qwen6nfe
pip install -r qwen6nfe/requirements.txt
python qwen6nfe/inference.py --prompt "A mountain lake at sunrise" --output mountain.png

The final command downloads the checkpoint and the necessary official base-model components automatically, then runs six-step inference. No training repository or custom CUDA extension is required. Downloads are cached for later calls.

Image editing with your own reference:

python qwen6nfe/inference.py --image input.png --prompt "Change the background to a sunny garden" --output edited.png

Your own prompt / reference:

python qwen6nfe/inference.py --prompt "A small cabin beside a mountain lake at sunrise" --width 2048 --height 2048 --seed 42 --output result.png
python qwen6nfe/inference.py --image input.png --prompt "Change the scene to late autumn, preserving the stone bridge" --output edited.png

Add --cpu-offload to reduce GPU memory use at the cost of speed; sufficient CPU RAM is still needed. High-resolution editing with reference images uses more memory than text-to-image.

To download the whole repository for local use:

hf download Sutoonq/Qwen-Image-2.1-6NFE-Distill --local-dir qwen6nfe
python qwen6nfe/inference.py --model qwen6nfe --prompt "A mountain lake at sunrise"

The official base components are still required and downloaded on first use.

Repository layout and download statistics

The root config.json is an exact copy of transformer/config.json, describing QwenImage21Transformer2DModel. The weights remain in transformer/, so component loading still requires subfolder="transformer". The root copy is included in the quick-start download for configuration inspection and download-statistics troubleshooting; it does not make this a self-contained pipeline.

model_index.json describes the pipeline components, pinned base model, and six-step inference settings. This is a transformer-only checkpoint, not a self-contained Diffusers pipeline: use the runner or the integration example below to obtain the remaining components from the official base model. The runner reads _inference_settings from this file and caches it alongside normal Hugging Face downloads. Older local exports without the file retain the same six-step defaults.

Hugging Face counts Diffusers downloads using the root model_index.json (or eligible root-level weight files). Whole-repository downloads and the updated runner include this file. Downloading only transformer/ bypasses that counting path. The root config.json has also been added as a diagnostic experiment: the documented Diffusers counting filter lists model_index.json and eligible root-level weights, whereas the default filter for repositories without a library includes config.json. Whether the root configuration changes this repository's observed count is not yet confirmed. Statistics are maintained by Hugging Face. See the official download statistics documentation.

Inference settings

  • 6 model evaluations, Euler flow sampling.
  • Conditional-only inference, true_cfg_scale=1.0 (no external CFG branch).
  • Preserve the base scheduler's resolution-dependent dynamic shift, but set shift_terminal=0.4. The sixth evaluated sigma is 0.4 and the final Euler update maps to zero. The runner applies this automatically.
  • Default output resolution: approximately 2048 pixels per side; explicit width and height can be supplied and are aligned by the official pipeline.
  • The pinned Diffusers revision and Transformers version are in requirements.txt.

Minimal integration into an existing Diffusers application:

import json
import torch
from huggingface_hub import hf_hub_download
from diffusers import (
    QwenImage21Pipeline, QwenImage21Transformer2DModel,
    FlowMatchEulerDiscreteScheduler,
)

with open(hf_hub_download(
    "Sutoonq/Qwen-Image-2.1-6NFE-Distill", "model_index.json",
), encoding="utf-8") as f:
    settings = json.load(f)["_inference_settings"]

transformer = QwenImage21Transformer2DModel.from_pretrained(
    "Sutoonq/Qwen-Image-2.1-6NFE-Distill", subfolder="transformer",
    torch_dtype=torch.bfloat16,
)
pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1",
    revision="b3179ad355be050328e483a9dfdd9e60cd62adfa",
    transformer=transformer, torch_dtype=torch.bfloat16,
).to("cuda")
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(
    pipe.scheduler.config, shift_terminal=settings["shift_terminal"],
)
image = pipe(
    prompt="A mountain lake at sunrise",
    num_inference_steps=settings["num_inference_steps"],
    true_cfg_scale=settings["true_cfg_scale"], width=2048, height=2048,
    output_resolution=2048,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("result.png")

Selected examples

Text-to-image

Text-to-image gallery

Prompts and results

Each column shows the prompt above its generated image. The Prompt Enhancement entries below summarize the longer prompts used for generation.

Short Prompt

设计一张精美的中文城市书展海报。标题必须准确写“让阅读照亮城市”,副标题“2026 秋季独立书展”,下方三列清晰排版:“时间 10月16日至18日”“地点 江畔艺术中心”“开放时间 10:00—20:00”。中央是一座由翻开的书页组成的立体城市,暖黄色窗光,深蓝夜空,纸张纹理,极简但有丰富细节,大小字体层级分明,所有文字完整清晰。

Text-to-image result

Short Prompt

中国茶文化展览海报,准确标题“山水之间,一盏清欢”,小字“东方茶事特展”,底部“春茶·器物·生活美学”。一只半透明青瓷茶杯放在湿润石台上,杯中倒映层叠群山,蒸汽形成细小白鹤,右侧有竹影,左侧竖排四行:“观山”“听雨”“品茗”“静心”。现代博物馆视觉设计,米白宣纸背景,墨绿和金色点缀,字形正确,构图留白。

Text-to-image result

Short Prompt

精致音乐节宣传海报,标题“山谷回声”,副标题“自然与声音的相遇”,日期“2026年8月22—23日”。中心是一把木吉他,琴身内是一片迷你森林和溪流,三名很小的音乐家站在琴桥上,萤火虫绕着琴弦。底部排版“森林舞台 / 星空露营 / 日落现场”。复古丝网印刷风,青绿、奶油白与橘红三色,细腻纸纹,准确中文。

Text-to-image result

Short Prompt

A photorealistic backstage fashion photograph with four distinct people: a model in a silver gown seated at center, a makeup artist standing to her left applying lipstick, a hairstylist behind her adjusting an elaborate braid, and a photographer reflected only in the mirror on the right. Correct mirror perspective, makeup lights, garment racks, naturally posed fingers, subtle cinematic grain, neutral skin tones, busy but readable composition.

Text-to-image result

Short Prompt

A detailed Japanese animation film still of six friends sheltering under a small countryside station roof during summer rain. Two share a red umbrella, one holds a sleeping cat, one checks a folded paper map, and two exchange a steaming drink. Wet bicycles stand at the left, hydrangeas bloom at the right, distant fields shimmer through rain. Each person visually distinct, convincing interaction, warm-cool light contrast, no lettering.

Text-to-image result

Short Prompt

A dramatic food photograph of freshly baked sourdough being cut on a walnut board. A serrated knife is halfway through the loaf, crisp crumbs scatter naturally, steam emerges from the open crumb, a ceramic dish of olive oil and rosemary sits behind it. Strong side light, flour dust in the air, accurately textured crust and porous interior, warm cinematic palette, no hands or lettering.

Text-to-image result

Prompt Enhancement

海岸暮色中的动漫电影海报,搭配西班牙语标题、奖项栏和上映信息。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

霓虹音乐活动海报,融合动漫主角、多格演出画面与日文活动文案。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

浅蓝鞋履广告:巨型产品、三位模特和分层宣传文字。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

画室人物摄影与彩色手绘涂鸦结合,营造专注创作氛围。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

犯罪题材游戏封面式拼贴,以角色、城市与大标题构成分区画面。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

热带夜景音乐封面,舞者剪影叠加城市灯光与霓虹标题。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

复古纸张上的男性肖像,结合钢笔线条、局部研究和手写笔记。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

高尔夫人物三联画,包含准备、细节特写与击球动作。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

咖啡和冰淇淋近景,辅以可爱插画与手写装饰。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

白色背景的舞蹈动作分格表,保持人物和服装一致。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

九宫格展示女性肖像从构形到完成的素描过程。

Original prompt: YouArt

Text-to-image result

Prompt Enhancement

复古香农信息论海报,结合人物、通信示意与历史时间线。

Original prompt: YouArt

Text-to-image result

Image editing

Reference images on the left; edited results on the right.

Image editing gallery

License

The Qwen Research License and official base-model terms apply.

Downloads last month
156
Safetensors
Model size
7B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sutoonq/Qwen-Image-2.1-6NFE-Distill

Finetuned
(61)
this model