Instructions to use kingjones777/Qwen-Image-2.1-ROCm-gfx1151 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use kingjones777/Qwen-Image-2.1-ROCm-gfx1151 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("kingjones777/Qwen-Image-2.1-ROCm-gfx1151", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Qwen-Image-2.1 on gfx1151: cold vs warm, MIOpen kernel-db finding
Browse files
README.md
ADDED
|
@@ -0,0 +1,140 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen-Image-2.1
|
| 4 |
+
pipeline_tag: text-to-image
|
| 5 |
+
library_name: diffusers
|
| 6 |
+
tags:
|
| 7 |
+
- qwen-image
|
| 8 |
+
- rocm
|
| 9 |
+
- amd
|
| 10 |
+
- gfx1151
|
| 11 |
+
- strix-halo
|
| 12 |
+
- benchmark
|
| 13 |
+
- documentation
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# Qwen-Image-2.1 on AMD Strix Halo (gfx1151, ROCm 7.13)
|
| 17 |
+
|
| 18 |
+
**This repository contains no weights.** It records what it takes to run
|
| 19 |
+
[`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) on an AMD Ryzen AI Max+
|
| 20 |
+
395 (Radeon 8060S, gfx1151) under ROCm 7.13, and the measured cost of doing so.
|
| 21 |
+
|
| 22 |
+
It runs. The first image takes **19.8 minutes** and the second takes **6.8 minutes**, and the
|
| 23 |
+
difference is not the model.
|
| 24 |
+
|
| 25 |
+
## Measured
|
| 26 |
+
|
| 27 |
+
Radeon 8060S / gfx1151, ROCm 7.13, torch 2.10.0 (hip 7.13.99004), diffusers 0.41.0.dev0,
|
| 28 |
+
bf16, 1024x1024, 24 steps, box otherwise idle. Two runs, same resolution and step count,
|
| 29 |
+
different prompt and seed.
|
| 30 |
+
|
| 31 |
+
| phase | run 1 (cold) | run 2 (warm) |
|
| 32 |
+
| --- | ---: | ---: |
|
| 33 |
+
| pipeline load to GPU | 7.3 s | — |
|
| 34 |
+
| sampler, 24 steps | 365.0 s (15.22 s/it) | 390.4 s (16.28 s/it) |
|
| 35 |
+
| **VAE decode** | **825.6 s** | **20.5 s** |
|
| 36 |
+
| total (`GEN_TIME_S`) | **1190.6 s** | **410.9 s** |
|
| 37 |
+
| peak GPU allocated | 40.55 GB | 36.89 GB |
|
| 38 |
+
| peak GPU reserved | 42.82 GB | 38.60 GB |
|
| 39 |
+
|
| 40 |
+
**The VAE decode got 40x faster on the second run.** Nothing about the model or the code
|
| 41 |
+
changed between them.
|
| 42 |
+
|
| 43 |
+
## ⛔ Why: ROCm ships no MIOpen kernel database for gfx1151
|
| 44 |
+
|
| 45 |
+
```
|
| 46 |
+
MIOpen(HIP): Warning [ParseAndLoadDb] File is unreadable:
|
| 47 |
+
"/opt/rocm/share/miopen/db/gfx1151_20.HIP.fdb.txt"
|
| 48 |
+
```
|
| 49 |
+
|
| 50 |
+
MIOpen ships pretuned kernel databases per GPU architecture. ROCm 7.13 has none for gfx1151,
|
| 51 |
+
so every convolution shape is **auto-tuned at runtime the first time it is seen**. The
|
| 52 |
+
transformer is unaffected (it is GEMM work), which is why the sampler rate is stable across
|
| 53 |
+
both runs. The VAE is convolutional, so it pays the entire tuning bill on run 1.
|
| 54 |
+
|
| 55 |
+
The tuned results persist in `~/.cache/miopen`. Evidence that this is the mechanism: the
|
| 56 |
+
cache was **221,184 bytes after run 1 and 221,184 bytes after run 2** — byte-identical, zero
|
| 57 |
+
new tuning, and the decode collapsed from 825.6 s to 20.5 s.
|
| 58 |
+
|
| 59 |
+
### It looks exactly like a hang
|
| 60 |
+
|
| 61 |
+
For ~13 minutes there is no log output, one host thread sits at 100%, and the process appears
|
| 62 |
+
stuck. It is not. Distinguish them by:
|
| 63 |
+
|
| 64 |
+
| signal | tuning | genuinely hung |
|
| 65 |
+
| --- | --- | --- |
|
| 66 |
+
| `/sys/class/drm/card0/device/gpu_busy_percent` | **94-100%** | ~0% |
|
| 67 |
+
| `~/.cache/miopen` size over a 20 s window | **growing** | static |
|
| 68 |
+
| process state | `Rl` | `D` / `S` |
|
| 69 |
+
|
| 70 |
+
⭐ **Do not quote a first-run timing as this model's speed on this hardware.** Warm the cache,
|
| 71 |
+
then measure. If a first run must be quick, `MIOPEN_FIND_MODE=FAST` shortens the search at
|
| 72 |
+
some cost in kernel quality — unset it for the run you actually report.
|
| 73 |
+
|
| 74 |
+
## The trap that silently costs you the GPU
|
| 75 |
+
|
| 76 |
+
This machine's **system** `python3` already had a working ROCm PyTorch
|
| 77 |
+
(`torch 2.10.0`, `torch.version.hip 7.13.99004`, `torch.cuda.is_available() True`), while
|
| 78 |
+
other virtualenvs on the box carried **CPU-only** torch. Letting pip resolve `torch` freshly,
|
| 79 |
+
or reusing the wrong venv, runs the entire 33 GB pipeline on CPU without ever erroring.
|
| 80 |
+
|
| 81 |
+
Build the venv so it inherits the working install, and check afterwards:
|
| 82 |
+
|
| 83 |
+
```bash
|
| 84 |
+
python3 -m venv --system-site-packages ~/build/qwen-image/venv
|
| 85 |
+
~/build/qwen-image/venv/bin/python -c \
|
| 86 |
+
"import torch; assert torch.cuda.is_available(); print(torch.__version__, torch.version.hip)"
|
| 87 |
+
```
|
| 88 |
+
|
| 89 |
+
⭐ Make the generator **refuse to run on CPU** rather than fall back — that turns a silent
|
| 90 |
+
30x slowdown into an immediate, obvious failure:
|
| 91 |
+
|
| 92 |
+
```python
|
| 93 |
+
if not torch.cuda.is_available():
|
| 94 |
+
print("FATAL: HIP not available - refusing to run on CPU"); sys.exit(2)
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
## Dependencies
|
| 98 |
+
|
| 99 |
+
`QwenImage21Pipeline` is newer than any diffusers release on PyPI — it needs
|
| 100 |
+
**diffusers >= 0.37.0.dev0**. Installed here from GitHub main via the archive zip (the build
|
| 101 |
+
box has no `git`):
|
| 102 |
+
|
| 103 |
+
```bash
|
| 104 |
+
pip install "diffusers @ https://github.com/huggingface/diffusers/archive/refs/heads/main.zip"
|
| 105 |
+
pip install "transformers>=5.17" accelerate safetensors
|
| 106 |
+
```
|
| 107 |
+
|
| 108 |
+
The text encoder is `Qwen3VLForConditionalGeneration`, which is why transformers 5.17+ is
|
| 109 |
+
required. Component sizes: text_encoder 17.53 GB, transformer 14.23 GB
|
| 110 |
+
(`QwenImage21Transformer2DModel`, 32 layers, 32 heads x 128), vae 1.35 GB — **33.1 GB** total
|
| 111 |
+
in bf16, all of which is resident on the GPU during generation.
|
| 112 |
+
|
| 113 |
+
## Files
|
| 114 |
+
|
| 115 |
+
| file | what it is |
|
| 116 |
+
| --- | --- |
|
| 117 |
+
| `README.md` | this document |
|
| 118 |
+
| `gen.py` | the generation script used for both runs (GPU-or-die, prints receipts) |
|
| 119 |
+
|
| 120 |
+
## Reproduction
|
| 121 |
+
|
| 122 |
+
```bash
|
| 123 |
+
python3 -m venv --system-site-packages ~/build/qwen-image/venv
|
| 124 |
+
. ~/build/qwen-image/venv/bin/activate
|
| 125 |
+
pip install "diffusers @ https://github.com/huggingface/diffusers/archive/refs/heads/main.zip" \
|
| 126 |
+
"transformers>=5.17" accelerate safetensors
|
| 127 |
+
|
| 128 |
+
HF_HUB_DISABLE_XET=1 hf download Qwen/Qwen-Image-2.1 --local-dir ~/models/qwen-image-2.1
|
| 129 |
+
|
| 130 |
+
python gen.py --model ~/models/qwen-image-2.1 --out out/a.png \
|
| 131 |
+
--prompt "A photorealistic red-tailed hawk perched on a saguaro cactus at golden hour" \
|
| 132 |
+
--steps 24 --width 1024 --height 1024 --seed 42
|
| 133 |
+
```
|
| 134 |
+
|
| 135 |
+
Run it twice. The second run is the honest number.
|
| 136 |
+
|
| 137 |
+
## Licence
|
| 138 |
+
|
| 139 |
+
Apache-2.0. Weights belong to [Qwen](https://huggingface.co/Qwen/Qwen-Image-2.1) under their
|
| 140 |
+
own licence; this repository contains only measurements and a script.
|