kingjones777 commited on
Commit
f8a0ca1
·
verified ·
1 Parent(s): 49808e1

Qwen-Image-2.1 on gfx1151: cold vs warm, MIOpen kernel-db finding

Browse files
Files changed (1) hide show
  1. README.md +42 -6
README.md CHANGED
@@ -20,7 +20,8 @@ tags:
20
  395 (Radeon 8060S, gfx1151) under ROCm 7.13, and the measured cost of doing so.
21
 
22
  It runs. The first image takes **19.8 minutes** and the second takes **6.8 minutes**, and the
23
- difference is not the model.
 
24
 
25
  ## Measured
26
 
@@ -40,17 +41,33 @@ different prompt and seed.
40
  **The VAE decode got 40x faster on the second run.** Nothing about the model or the code
41
  changed between them.
42
 
43
- ## ⛔ Why: ROCm ships no MIOpen kernel database for gfx1151
44
 
45
  ```
46
  MIOpen(HIP): Warning [ParseAndLoadDb] File is unreadable:
47
  "/opt/rocm/share/miopen/db/gfx1151_20.HIP.fdb.txt"
48
  ```
49
 
50
- MIOpen ships pretuned kernel databases per GPU architecture. ROCm 7.13 has none for gfx1151,
51
- so every convolution shape is **auto-tuned at runtime the first time it is seen**. The
52
- transformer is unaffected (it is GEMM work), which is why the sampler rate is stable across
53
- both runs. The VAE is convolutional, so it pays the entire tuning bill on run 1.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
54
 
55
  The tuned results persist in `~/.cache/miopen`. Evidence that this is the mechanism: the
56
  cache was **221,184 bytes after run 1 and 221,184 bytes after run 2** — byte-identical, zero
@@ -71,6 +88,25 @@ stuck. It is not. Distinguish them by:
71
  then measure. If a first run must be quick, `MIOPEN_FIND_MODE=FAST` shortens the search at
72
  some cost in kernel quality — unset it for the run you actually report.
73
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
74
  ## The trap that silently costs you the GPU
75
 
76
  This machine's **system** `python3` already had a working ROCm PyTorch
 
20
  395 (Radeon 8060S, gfx1151) under ROCm 7.13, and the measured cost of doing so.
21
 
22
  It runs. The first image takes **19.8 minutes** and the second takes **6.8 minutes**, and the
23
+ difference is not the model — it is MIOpen tuning convolution kernels once, on a ROCm build
24
+ that has no pretuned database for this GPU.
25
 
26
  ## Measured
27
 
 
41
  **The VAE decode got 40x faster on the second run.** Nothing about the model or the code
42
  changed between them.
43
 
44
+ ## ⛔ Why: no MIOpen *tuning database* for gfx1151 on this ROCm build
45
 
46
  ```
47
  MIOpen(HIP): Warning [ParseAndLoadDb] File is unreadable:
48
  "/opt/rocm/share/miopen/db/gfx1151_20.HIP.fdb.txt"
49
  ```
50
 
51
+ MIOpen ships pretuned performance databases per GPU architecture. On **this install
52
+ (ROCm 7.13)** there is none for gfx1151: `/opt/rocm/share/miopen/db/` holds 81 files covering
53
+ gfx908, gfx90a and gfx942, and **zero** matching `gfx1151`. So every convolution shape is
54
+ **auto-tuned at runtime the first time it is seen**. The transformer is unaffected (it is
55
+ GEMM work), which is why the sampler rate is stable across both runs. The VAE is
56
+ convolutional, so it pays the entire tuning bill on run 1.
57
+
58
+ ⭐ **Two separate things, and only one is missing here.** The CK grouped-convolution *kernel
59
+ library* `libMIOpenCKGroupedConv_gfx1151.so` **is present** on this system. Only the *tuning
60
+ database* is absent. That distinction matters: the related upstream report
61
+ [ROCm/TheRock#5105](https://github.com/ROCm/TheRock/issues/5105) describes a worse case where
62
+ **both** were missing and convolutions fell back to the `GemmFwdRest` solver, making an
63
+ inference workload 3-5x slower than CPU. If you see
64
+ `CK grouped conv library not found for device gfx1151` in addition to the fdb warning, you
65
+ have that bug, not this one.
66
+
67
+ ⚠️ **Version scope.** These numbers are ROCm 7.13. AMD has since moved gfx1151 packaging
68
+ forward (ROCm 10.0 publishes dedicated `device-gfx1151` payloads). **We have not tested
69
+ ROCm 10 on this hardware**, so do not read this page as a claim about current ROCm — it is a
70
+ measurement of one released version, and the *method* below is what generalises.
71
 
72
  The tuned results persist in `~/.cache/miopen`. Evidence that this is the mechanism: the
73
  cache was **221,184 bytes after run 1 and 221,184 bytes after run 2** — byte-identical, zero
 
88
  then measure. If a first run must be quick, `MIOPEN_FIND_MODE=FAST` shortens the search at
89
  some cost in kernel quality — unset it for the run you actually report.
90
 
91
+ ### You can build the tuning database yourself
92
+
93
+ A missing *system* performance database is not a dead end — MIOpen consults a user PerfDb
94
+ that overrides it, and AMD documents generating one
95
+ ([tuning performance databases](https://rocm.docs.amd.com/projects/MIOpen/en/develop/conceptual/tuningdb.html)).
96
+ Exercise your real shapes once with search enabled:
97
+
98
+ ```bash
99
+ export MIOPEN_USER_DB_PATH="$HOME/.config/miopen"
100
+ export MIOPEN_FIND_MODE=NORMAL
101
+ export MIOPEN_FIND_ENFORCE=SEARCH_DB_UPDATE
102
+ python gen.py ... # run the resolutions you actually use
103
+ unset MIOPEN_FIND_MODE MIOPEN_FIND_ENFORCE
104
+ ```
105
+
106
+ Subsequent runs read the tuned entries. This is what the warm run above is doing implicitly —
107
+ explicit tuning just lets you front-load it deliberately, per shape, instead of paying it
108
+ inside a user-facing request.
109
+
110
  ## The trap that silently costs you the GPU
111
 
112
  This machine's **system** `python3` already had a working ROCm PyTorch