djdeniro commited on
Commit
925e115
Β·
verified Β·
1 Parent(s): 290852f

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +95 -0
README.md ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ base_model: zai-org/GLM-5.3-Flash
4
+ tags:
5
+ - glm
6
+ - mla
7
+ - linear-attention
8
+ - moe
9
+ - multimodal
10
+ - rocm
11
+ - rdna4
12
+ - gfx1201
13
+ - rfa
14
+ - rfi
15
+ - amd
16
+ - text-generation
17
+ ---
18
+
19
+ # GLM-5.3-Flash RFA-RFI8 β€” FOR INFERENCE ON 8Γ— R9700 (RDNA4/gfx1201)
20
+
21
+ A **self-quantized derivative** of
22
+ **[zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash)** (321B MoE, β‰ˆ18B active),
23
+ quantized with a **RFA + RFI8 composite** and validated for serving on
24
+ **8Γ— AMD Radeon R9700 (gfx1201 / RDNA4)**.
25
+
26
+ > ⚠️ **Not compatible with stock vLLM.** This checkpoint uses the `rfi` composite quantizer and the
27
+ > `Glm5NextForConditionalGeneration` architecture on the RDNA4 path, which requires the patched
28
+ > vLLM + overlay in **[djdeniro/GLM-5.3-Flash-rocm-r9700](https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700)**.
29
+
30
+ ---
31
+
32
+ ## Model
33
+
34
+ - **Base:** [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash) β€” MIT license.
35
+ - **Architecture:** 321B MoE, β‰ˆ18B active; **45 layers = 34 KDA (linear attention) + 11 DSA
36
+ (sparse-MLA)**; mHC hidden-state compression; **native multimodal** (image + video); 1 nextn MTP
37
+ layer; 288 routed experts (top-8) + 1 shared expert.
38
+ - **Paper:** [arXiv:2602.15763](https://arxiv.org/abs/2602.15763).
39
+
40
+ ## Quantization
41
+
42
+ | Scheme | Applied to | bpw |
43
+ |--------|-----------|-----|
44
+ | **RFA** | MoE routed experts | 4.5 |
45
+ | **RFI8** (int8 W8A8) | attention / shared / dense linears (structural) | 8.25 |
46
+ | **BF16** | dense copies + MTP layer | 16 |
47
+
48
+ Total **β‰ˆ197.8 GB** (25 safetensors shards). Full recipe (archspec, source patches, kda-remap,
49
+ run scripts) is in the code repo: **[djdeniro/GLM-5.3-Flash-rocm-r9700](https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700)** (`quant/`).
50
+
51
+ ## Quick start (8Γ— R9700)
52
+
53
+ ```bash
54
+ git clone https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700 overlay
55
+ huggingface-cli download djdeniro/GLM-5.3-Flash-RFA-RFI8-8xR9700 --local-dir ./models
56
+
57
+ docker run --rm --tty --ipc=host --shm-size=128g \
58
+ --device /dev/kfd:/dev/kfd --device /dev/dri:/dev/dri \
59
+ -v "$PWD/models":/models:ro -v "$PWD/overlay":/overlay:ro \
60
+ --entrypoint bash tcclaviger/vllm:latest \
61
+ -c "/overlay/apply_overlay.sh && exec vllm serve /models \
62
+ --served-model-name glm53-flash --trust-remote-code --quantization rfi \
63
+ --tensor-parallel-size 8 --gpu-memory-utilization 0.95 \
64
+ --max-model-len 190080 --max-num-seqs 4 --kv-cache-dtype auto"
65
+ ```
66
+
67
+ Use `overlay/run-glm53.sh` (or `denet-large.sh` via llama-swap) for the full production command.
68
+
69
+ ## Performance (8Γ— R9700, gfx1201)
70
+
71
+ - **Decode:** β‰ˆ **34–37 t/s** at bs=1 (FULL cudagraph).
72
+ - **TTFT:** β‰ˆ **0.13–0.6 s** (prompt-dependent, `reasoning_effort="low"`).
73
+ - **Context:** **190k** tokens with bf16 KV.
74
+
75
+ ## Multimodal
76
+
77
+ Images are resized preserving aspect ratio: **min 384Γ—384** (upscale) / **max 1280Γ—1280**
78
+ (downscale). The processor/video tower is the native Glm5Next multimodal path.
79
+
80
+ ## Known limitations
81
+
82
+ - **MTP is OFF** β€” the nextn drafter is blocked by vLLM's kv-cache-group assertion.
83
+ - **Do NOT enable fp8 KV** (`--kv-cache-dtype fp8` + `--calculate-kv-scales`): runtime scale
84
+ calibration runs through an unwarmed KDA recurrent state on the profile dummy-run, producing
85
+ wrong `_k_scale` β†’ hard output looping (upstream vLLM issue **#37554**). The checkpoint ships no
86
+ static KV scales β€” serve with **bf16 KV** (`--kv-cache-dtype auto`).
87
+ - **Chat:** send `reasoning_effort="low"` (the GLM-5.3 chat template defaults to Reasoning Effort
88
+ Max and over-thinks on long generations).
89
+
90
+ ## License & attribution
91
+
92
+ - **License:** MIT.
93
+ - **Original model:** [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash) β€” Β© Z.ai
94
+ (zai-org), MIT license.
95
+ - **Paper:** [arXiv:2602.15763](https://arxiv.org/abs/2602.15763).