Instructions to use kingjones777/Qwen-Image-2.1-ROCm-gfx1151 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use kingjones777/Qwen-Image-2.1-ROCm-gfx1151 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("kingjones777/Qwen-Image-2.1-ROCm-gfx1151", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Qwen-Image-2.1 on gfx1151: cold vs warm, MIOpen kernel-db finding
Browse files
README.md
CHANGED
|
@@ -20,7 +20,8 @@ tags:
|
|
| 20 |
395 (Radeon 8060S, gfx1151) under ROCm 7.13, and the measured cost of doing so.
|
| 21 |
|
| 22 |
It runs. The first image takes **19.8 minutes** and the second takes **6.8 minutes**, and the
|
| 23 |
-
difference is not the model
|
|
|
|
| 24 |
|
| 25 |
## Measured
|
| 26 |
|
|
@@ -40,17 +41,33 @@ different prompt and seed.
|
|
| 40 |
**The VAE decode got 40x faster on the second run.** Nothing about the model or the code
|
| 41 |
changed between them.
|
| 42 |
|
| 43 |
-
## ⛔ Why:
|
| 44 |
|
| 45 |
```
|
| 46 |
MIOpen(HIP): Warning [ParseAndLoadDb] File is unreadable:
|
| 47 |
"/opt/rocm/share/miopen/db/gfx1151_20.HIP.fdb.txt"
|
| 48 |
```
|
| 49 |
|
| 50 |
-
MIOpen ships pretuned
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 54 |
|
| 55 |
The tuned results persist in `~/.cache/miopen`. Evidence that this is the mechanism: the
|
| 56 |
cache was **221,184 bytes after run 1 and 221,184 bytes after run 2** — byte-identical, zero
|
|
@@ -71,6 +88,25 @@ stuck. It is not. Distinguish them by:
|
|
| 71 |
then measure. If a first run must be quick, `MIOPEN_FIND_MODE=FAST` shortens the search at
|
| 72 |
some cost in kernel quality — unset it for the run you actually report.
|
| 73 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
## The trap that silently costs you the GPU
|
| 75 |
|
| 76 |
This machine's **system** `python3` already had a working ROCm PyTorch
|
|
|
|
| 20 |
395 (Radeon 8060S, gfx1151) under ROCm 7.13, and the measured cost of doing so.
|
| 21 |
|
| 22 |
It runs. The first image takes **19.8 minutes** and the second takes **6.8 minutes**, and the
|
| 23 |
+
difference is not the model — it is MIOpen tuning convolution kernels once, on a ROCm build
|
| 24 |
+
that has no pretuned database for this GPU.
|
| 25 |
|
| 26 |
## Measured
|
| 27 |
|
|
|
|
| 41 |
**The VAE decode got 40x faster on the second run.** Nothing about the model or the code
|
| 42 |
changed between them.
|
| 43 |
|
| 44 |
+
## ⛔ Why: no MIOpen *tuning database* for gfx1151 on this ROCm build
|
| 45 |
|
| 46 |
```
|
| 47 |
MIOpen(HIP): Warning [ParseAndLoadDb] File is unreadable:
|
| 48 |
"/opt/rocm/share/miopen/db/gfx1151_20.HIP.fdb.txt"
|
| 49 |
```
|
| 50 |
|
| 51 |
+
MIOpen ships pretuned performance databases per GPU architecture. On **this install
|
| 52 |
+
(ROCm 7.13)** there is none for gfx1151: `/opt/rocm/share/miopen/db/` holds 81 files covering
|
| 53 |
+
gfx908, gfx90a and gfx942, and **zero** matching `gfx1151`. So every convolution shape is
|
| 54 |
+
**auto-tuned at runtime the first time it is seen**. The transformer is unaffected (it is
|
| 55 |
+
GEMM work), which is why the sampler rate is stable across both runs. The VAE is
|
| 56 |
+
convolutional, so it pays the entire tuning bill on run 1.
|
| 57 |
+
|
| 58 |
+
⭐ **Two separate things, and only one is missing here.** The CK grouped-convolution *kernel
|
| 59 |
+
library* `libMIOpenCKGroupedConv_gfx1151.so` **is present** on this system. Only the *tuning
|
| 60 |
+
database* is absent. That distinction matters: the related upstream report
|
| 61 |
+
[ROCm/TheRock#5105](https://github.com/ROCm/TheRock/issues/5105) describes a worse case where
|
| 62 |
+
**both** were missing and convolutions fell back to the `GemmFwdRest` solver, making an
|
| 63 |
+
inference workload 3-5x slower than CPU. If you see
|
| 64 |
+
`CK grouped conv library not found for device gfx1151` in addition to the fdb warning, you
|
| 65 |
+
have that bug, not this one.
|
| 66 |
+
|
| 67 |
+
⚠️ **Version scope.** These numbers are ROCm 7.13. AMD has since moved gfx1151 packaging
|
| 68 |
+
forward (ROCm 10.0 publishes dedicated `device-gfx1151` payloads). **We have not tested
|
| 69 |
+
ROCm 10 on this hardware**, so do not read this page as a claim about current ROCm — it is a
|
| 70 |
+
measurement of one released version, and the *method* below is what generalises.
|
| 71 |
|
| 72 |
The tuned results persist in `~/.cache/miopen`. Evidence that this is the mechanism: the
|
| 73 |
cache was **221,184 bytes after run 1 and 221,184 bytes after run 2** — byte-identical, zero
|
|
|
|
| 88 |
then measure. If a first run must be quick, `MIOPEN_FIND_MODE=FAST` shortens the search at
|
| 89 |
some cost in kernel quality — unset it for the run you actually report.
|
| 90 |
|
| 91 |
+
### You can build the tuning database yourself
|
| 92 |
+
|
| 93 |
+
A missing *system* performance database is not a dead end — MIOpen consults a user PerfDb
|
| 94 |
+
that overrides it, and AMD documents generating one
|
| 95 |
+
([tuning performance databases](https://rocm.docs.amd.com/projects/MIOpen/en/develop/conceptual/tuningdb.html)).
|
| 96 |
+
Exercise your real shapes once with search enabled:
|
| 97 |
+
|
| 98 |
+
```bash
|
| 99 |
+
export MIOPEN_USER_DB_PATH="$HOME/.config/miopen"
|
| 100 |
+
export MIOPEN_FIND_MODE=NORMAL
|
| 101 |
+
export MIOPEN_FIND_ENFORCE=SEARCH_DB_UPDATE
|
| 102 |
+
python gen.py ... # run the resolutions you actually use
|
| 103 |
+
unset MIOPEN_FIND_MODE MIOPEN_FIND_ENFORCE
|
| 104 |
+
```
|
| 105 |
+
|
| 106 |
+
Subsequent runs read the tuned entries. This is what the warm run above is doing implicitly —
|
| 107 |
+
explicit tuning just lets you front-load it deliberately, per shape, instead of paying it
|
| 108 |
+
inside a user-facing request.
|
| 109 |
+
|
| 110 |
## The trap that silently costs you the GPU
|
| 111 |
|
| 112 |
This machine's **system** `python3` already had a working ROCm PyTorch
|