kingjones777 commited on
Commit
f18707b
·
verified ·
1 Parent(s): 3bff06c

Qwen-Image-2.1 on gfx1151: cold vs warm, MIOpen kernel-db finding

Browse files
Files changed (1) hide show
  1. README.md +140 -0
README.md ADDED
@@ -0,0 +1,140 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen-Image-2.1
4
+ pipeline_tag: text-to-image
5
+ library_name: diffusers
6
+ tags:
7
+ - qwen-image
8
+ - rocm
9
+ - amd
10
+ - gfx1151
11
+ - strix-halo
12
+ - benchmark
13
+ - documentation
14
+ ---
15
+
16
+ # Qwen-Image-2.1 on AMD Strix Halo (gfx1151, ROCm 7.13)
17
+
18
+ **This repository contains no weights.** It records what it takes to run
19
+ [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) on an AMD Ryzen AI Max+
20
+ 395 (Radeon 8060S, gfx1151) under ROCm 7.13, and the measured cost of doing so.
21
+
22
+ It runs. The first image takes **19.8 minutes** and the second takes **6.8 minutes**, and the
23
+ difference is not the model.
24
+
25
+ ## Measured
26
+
27
+ Radeon 8060S / gfx1151, ROCm 7.13, torch 2.10.0 (hip 7.13.99004), diffusers 0.41.0.dev0,
28
+ bf16, 1024x1024, 24 steps, box otherwise idle. Two runs, same resolution and step count,
29
+ different prompt and seed.
30
+
31
+ | phase | run 1 (cold) | run 2 (warm) |
32
+ | --- | ---: | ---: |
33
+ | pipeline load to GPU | 7.3 s | — |
34
+ | sampler, 24 steps | 365.0 s (15.22 s/it) | 390.4 s (16.28 s/it) |
35
+ | **VAE decode** | **825.6 s** | **20.5 s** |
36
+ | total (`GEN_TIME_S`) | **1190.6 s** | **410.9 s** |
37
+ | peak GPU allocated | 40.55 GB | 36.89 GB |
38
+ | peak GPU reserved | 42.82 GB | 38.60 GB |
39
+
40
+ **The VAE decode got 40x faster on the second run.** Nothing about the model or the code
41
+ changed between them.
42
+
43
+ ## ⛔ Why: ROCm ships no MIOpen kernel database for gfx1151
44
+
45
+ ```
46
+ MIOpen(HIP): Warning [ParseAndLoadDb] File is unreadable:
47
+ "/opt/rocm/share/miopen/db/gfx1151_20.HIP.fdb.txt"
48
+ ```
49
+
50
+ MIOpen ships pretuned kernel databases per GPU architecture. ROCm 7.13 has none for gfx1151,
51
+ so every convolution shape is **auto-tuned at runtime the first time it is seen**. The
52
+ transformer is unaffected (it is GEMM work), which is why the sampler rate is stable across
53
+ both runs. The VAE is convolutional, so it pays the entire tuning bill on run 1.
54
+
55
+ The tuned results persist in `~/.cache/miopen`. Evidence that this is the mechanism: the
56
+ cache was **221,184 bytes after run 1 and 221,184 bytes after run 2** — byte-identical, zero
57
+ new tuning, and the decode collapsed from 825.6 s to 20.5 s.
58
+
59
+ ### It looks exactly like a hang
60
+
61
+ For ~13 minutes there is no log output, one host thread sits at 100%, and the process appears
62
+ stuck. It is not. Distinguish them by:
63
+
64
+ | signal | tuning | genuinely hung |
65
+ | --- | --- | --- |
66
+ | `/sys/class/drm/card0/device/gpu_busy_percent` | **94-100%** | ~0% |
67
+ | `~/.cache/miopen` size over a 20 s window | **growing** | static |
68
+ | process state | `Rl` | `D` / `S` |
69
+
70
+ ⭐ **Do not quote a first-run timing as this model's speed on this hardware.** Warm the cache,
71
+ then measure. If a first run must be quick, `MIOPEN_FIND_MODE=FAST` shortens the search at
72
+ some cost in kernel quality — unset it for the run you actually report.
73
+
74
+ ## The trap that silently costs you the GPU
75
+
76
+ This machine's **system** `python3` already had a working ROCm PyTorch
77
+ (`torch 2.10.0`, `torch.version.hip 7.13.99004`, `torch.cuda.is_available() True`), while
78
+ other virtualenvs on the box carried **CPU-only** torch. Letting pip resolve `torch` freshly,
79
+ or reusing the wrong venv, runs the entire 33 GB pipeline on CPU without ever erroring.
80
+
81
+ Build the venv so it inherits the working install, and check afterwards:
82
+
83
+ ```bash
84
+ python3 -m venv --system-site-packages ~/build/qwen-image/venv
85
+ ~/build/qwen-image/venv/bin/python -c \
86
+ "import torch; assert torch.cuda.is_available(); print(torch.__version__, torch.version.hip)"
87
+ ```
88
+
89
+ ⭐ Make the generator **refuse to run on CPU** rather than fall back — that turns a silent
90
+ 30x slowdown into an immediate, obvious failure:
91
+
92
+ ```python
93
+ if not torch.cuda.is_available():
94
+ print("FATAL: HIP not available - refusing to run on CPU"); sys.exit(2)
95
+ ```
96
+
97
+ ## Dependencies
98
+
99
+ `QwenImage21Pipeline` is newer than any diffusers release on PyPI — it needs
100
+ **diffusers >= 0.37.0.dev0**. Installed here from GitHub main via the archive zip (the build
101
+ box has no `git`):
102
+
103
+ ```bash
104
+ pip install "diffusers @ https://github.com/huggingface/diffusers/archive/refs/heads/main.zip"
105
+ pip install "transformers>=5.17" accelerate safetensors
106
+ ```
107
+
108
+ The text encoder is `Qwen3VLForConditionalGeneration`, which is why transformers 5.17+ is
109
+ required. Component sizes: text_encoder 17.53 GB, transformer 14.23 GB
110
+ (`QwenImage21Transformer2DModel`, 32 layers, 32 heads x 128), vae 1.35 GB — **33.1 GB** total
111
+ in bf16, all of which is resident on the GPU during generation.
112
+
113
+ ## Files
114
+
115
+ | file | what it is |
116
+ | --- | --- |
117
+ | `README.md` | this document |
118
+ | `gen.py` | the generation script used for both runs (GPU-or-die, prints receipts) |
119
+
120
+ ## Reproduction
121
+
122
+ ```bash
123
+ python3 -m venv --system-site-packages ~/build/qwen-image/venv
124
+ . ~/build/qwen-image/venv/bin/activate
125
+ pip install "diffusers @ https://github.com/huggingface/diffusers/archive/refs/heads/main.zip" \
126
+ "transformers>=5.17" accelerate safetensors
127
+
128
+ HF_HUB_DISABLE_XET=1 hf download Qwen/Qwen-Image-2.1 --local-dir ~/models/qwen-image-2.1
129
+
130
+ python gen.py --model ~/models/qwen-image-2.1 --out out/a.png \
131
+ --prompt "A photorealistic red-tailed hawk perched on a saguaro cactus at golden hour" \
132
+ --steps 24 --width 1024 --height 1024 --seed 42
133
+ ```
134
+
135
+ Run it twice. The second run is the honest number.
136
+
137
+ ## Licence
138
+
139
+ Apache-2.0. Weights belong to [Qwen](https://huggingface.co/Qwen/Qwen-Image-2.1) under their
140
+ own licence; this repository contains only measurements and a script.