--- license: mit base_model: zai-org/GLM-5.3-Flash tags: - glm - mla - linear-attention - moe - multimodal - rocm - rdna4 - gfx1201 - rfa - rfi - amd - text-generation --- # GLM-5.3-Flash RFA-RFI8 — FOR INFERENCE ON 8× R9700 (RDNA4/gfx1201) A **self-quantized derivative** of **[zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash)** (321B MoE, ≈18B active), quantized with a **RFA + RFI8 composite** and validated for serving on **8× AMD Radeon R9700 (gfx1201 / RDNA4)**. > ⚠️ **Not compatible with stock vLLM.** This checkpoint uses the `rfi` composite quantizer and the > `Glm5NextForConditionalGeneration` architecture on the RDNA4 path, which requires the patched > vLLM + overlay in **[djdeniro/GLM-5.3-Flash-rocm-r9700](https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700)**. --- ## Model - **Base:** [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash) — MIT license. - **Architecture:** 321B MoE, ≈18B active; **45 layers = 34 KDA (linear attention) + 11 DSA (sparse-MLA)**; mHC hidden-state compression; **native multimodal** (image + video); 1 nextn MTP layer; 288 routed experts (top-8) + 1 shared expert. - **Paper:** [arXiv:2602.15763](https://arxiv.org/abs/2602.15763). ## Quantization | Scheme | Applied to | bpw | |--------|-----------|-----| | **RFA** | MoE routed experts | 4.5 | | **RFI8** (int8 W8A8) | attention / shared / dense linears (structural) | 8.25 | | **BF16** | dense copies + MTP layer | 16 | Total **≈197.8 GB** (25 safetensors shards). Full recipe (archspec, source patches, kda-remap, run scripts) is in the code repo: **[djdeniro/GLM-5.3-Flash-rocm-r9700](https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700)** (`quant/`). ## Quick start (8× R9700) ```bash git clone https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700 overlay huggingface-cli download djdeniro/GLM-5.3-Flash-RFA-RFI8-8xR9700 --local-dir ./models docker run --rm --tty --ipc=host --shm-size=128g \ --device /dev/kfd:/dev/kfd --device /dev/dri:/dev/dri \ -v "$PWD/models":/models:ro -v "$PWD/overlay":/overlay:ro \ --entrypoint bash tcclaviger/vllm:latest \ -c "/overlay/apply_overlay.sh && exec vllm serve /models \ --served-model-name glm53-flash --trust-remote-code --quantization rfi \ --tensor-parallel-size 8 --gpu-memory-utilization 0.95 \ --max-model-len 190080 --max-num-seqs 4 --kv-cache-dtype auto" ``` Use `overlay/run-glm53.sh` (or `denet-large.sh` via llama-swap) for the full production command. ## Performance (8× R9700, gfx1201) - **Decode:** ≈ **34–37 t/s** at bs=1 (FULL cudagraph). - **TTFT:** ≈ **0.13–0.6 s** (prompt-dependent, `reasoning_effort="low"`). - **Context:** **190k** tokens with bf16 KV. ## Multimodal Images are resized preserving aspect ratio: **min 384×384** (upscale) / **max 1280×1280** (downscale). The processor/video tower is the native Glm5Next multimodal path. ## Known limitations - **MTP is OFF** — the nextn drafter is blocked by vLLM's kv-cache-group assertion. - **Do NOT enable fp8 KV** (`--kv-cache-dtype fp8` + `--calculate-kv-scales`): runtime scale calibration runs through an unwarmed KDA recurrent state on the profile dummy-run, producing wrong `_k_scale` → hard output looping (upstream vLLM issue **#37554**). The checkpoint ships no static KV scales — serve with **bf16 KV** (`--kv-cache-dtype auto`). - **Chat:** send `reasoning_effort="low"` (the GLM-5.3 chat template defaults to Reasoning Effort Max and over-thinks on long generations). ## License & attribution - **License:** MIT. - **Original model:** [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash) — © Z.ai (zai-org), MIT license. - **Paper:** [arXiv:2602.15763](https://arxiv.org/abs/2602.15763).