djdeniro's picture
Upload README.md with huggingface_hub
925e115 verified
|
Raw History Blame
3.81 kB
metadata
license: mit
base_model: zai-org/GLM-5.3-Flash
tags:
  - glm
  - mla
  - linear-attention
  - moe
  - multimodal
  - rocm
  - rdna4
  - gfx1201
  - rfa
  - rfi
  - amd
  - text-generation

GLM-5.3-Flash RFA-RFI8 β€” FOR INFERENCE ON 8Γ— R9700 (RDNA4/gfx1201)

A self-quantized derivative of zai-org/GLM-5.3-Flash (321B MoE, β‰ˆ18B active), quantized with a RFA + RFI8 composite and validated for serving on 8Γ— AMD Radeon R9700 (gfx1201 / RDNA4).

⚠️ Not compatible with stock vLLM. This checkpoint uses the rfi composite quantizer and the Glm5NextForConditionalGeneration architecture on the RDNA4 path, which requires the patched vLLM + overlay in djdeniro/GLM-5.3-Flash-rocm-r9700.


Model

  • Base: zai-org/GLM-5.3-Flash β€” MIT license.
  • Architecture: 321B MoE, β‰ˆ18B active; 45 layers = 34 KDA (linear attention) + 11 DSA (sparse-MLA); mHC hidden-state compression; native multimodal (image + video); 1 nextn MTP layer; 288 routed experts (top-8) + 1 shared expert.
  • Paper: arXiv:2602.15763.

Quantization

Scheme Applied to bpw
RFA MoE routed experts 4.5
RFI8 (int8 W8A8) attention / shared / dense linears (structural) 8.25
BF16 dense copies + MTP layer 16

Total β‰ˆ197.8 GB (25 safetensors shards). Full recipe (archspec, source patches, kda-remap, run scripts) is in the code repo: djdeniro/GLM-5.3-Flash-rocm-r9700 (quant/).

Quick start (8Γ— R9700)

git clone https://huggingface.co/djdeniro/GLM-5.3-Flash-rocm-r9700 overlay
huggingface-cli download djdeniro/GLM-5.3-Flash-RFA-RFI8-8xR9700 --local-dir ./models

docker run --rm --tty --ipc=host --shm-size=128g \
  --device /dev/kfd:/dev/kfd --device /dev/dri:/dev/dri \
  -v "$PWD/models":/models:ro -v "$PWD/overlay":/overlay:ro \
  --entrypoint bash tcclaviger/vllm:latest \
  -c "/overlay/apply_overlay.sh && exec vllm serve /models \
      --served-model-name glm53-flash --trust-remote-code --quantization rfi \
      --tensor-parallel-size 8 --gpu-memory-utilization 0.95 \
      --max-model-len 190080 --max-num-seqs 4 --kv-cache-dtype auto"

Use overlay/run-glm53.sh (or denet-large.sh via llama-swap) for the full production command.

Performance (8Γ— R9700, gfx1201)

  • Decode: β‰ˆ 34–37 t/s at bs=1 (FULL cudagraph).
  • TTFT: β‰ˆ 0.13–0.6 s (prompt-dependent, reasoning_effort="low").
  • Context: 190k tokens with bf16 KV.

Multimodal

Images are resized preserving aspect ratio: min 384Γ—384 (upscale) / max 1280Γ—1280 (downscale). The processor/video tower is the native Glm5Next multimodal path.

Known limitations

  • MTP is OFF β€” the nextn drafter is blocked by vLLM's kv-cache-group assertion.
  • Do NOT enable fp8 KV (--kv-cache-dtype fp8 + --calculate-kv-scales): runtime scale calibration runs through an unwarmed KDA recurrent state on the profile dummy-run, producing wrong _k_scale β†’ hard output looping (upstream vLLM issue #37554). The checkpoint ships no static KV scales β€” serve with bf16 KV (--kv-cache-dtype auto).
  • Chat: send reasoning_effort="low" (the GLM-5.3 chat template defaults to Reasoning Effort Max and over-thinks on long generations).

License & attribution