brandonmusic commited on
Commit
1ae6d70
·
verified ·
1 Parent(s): 3224669

Document selectable MTP3 and DFlash2 profiles

Browse files

Adds directly runnable multimodal DFlash2, language-only DFlash2, and language-only built-in MTP3 options with the measured KV-token capacity difference.

README.md CHANGED
@@ -19,9 +19,23 @@ tags:
19
  # GLM-5.3-Flash TR3 4bpw — current SM120 runtime
20
 
21
  This is the uniform-K4 EXL3/TR3 routed-expert checkpoint for GLM-5.3-Flash.
22
- The current daily-driver runtime is v84: TP2/EP2/DCP2, calibrated NVFP4 MLA
23
- KV, DFlash2-7, CUDA graphs, and working image input on two SM120 GPUs. It is a
24
- custom vLLM/B12X build and is not compatible with stock upstream vLLM.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
 
26
  ## Run the current image
27
 
@@ -91,6 +105,20 @@ GPU_DEVICES=0,1 \
91
  ./serve-glm53-sm120-tp2-language-only.sh
92
  ```
93
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
94
  The language-only alias points to the same tested v84 code digest:
95
 
96
  ```text
@@ -109,7 +137,10 @@ Measured capacity on the same two 96 GB GPUs at 300 W each:
109
  Turning vision off raises the DFlash KV token pool by 42.6%. The much larger
110
  7.45x language-only gain comes from using the built-in MTP head instead of
111
  keeping the external DFlash2-7 model resident. It is not a vision-only gain.
112
- The language-only profile does not accept image inputs.
 
 
 
113
 
114
  ## Current measured results
115
 
 
19
  # GLM-5.3-Flash TR3 4bpw — current SM120 runtime
20
 
21
  This is the uniform-K4 EXL3/TR3 routed-expert checkpoint for GLM-5.3-Flash.
22
+ The current v84 runtime supports three explicit TP2/EP2/DCP2 profiles on two
23
+ SM120 GPUs: multimodal DFlash2, language-only DFlash2, and language-only MTP3.
24
+ All use calibrated NVFP4 MLA KV and CUDA graphs. This is a custom vLLM/B12X
25
+ build and is not compatible with stock upstream vLLM.
26
+
27
+ ## Pick a serving profile
28
+
29
+ | Goal | Launcher | Extra checkpoint | Measured KV tokens |
30
+ |---|---|---|---:|
31
+ | Images plus fastest measured C1 decode | `compose.sm120-tp2.yaml` | DFlash2-7 | 129,473 |
32
+ | Text-only DFlash2 decode | `compose.sm120-tp2-language-only-dflash2.yaml` | DFlash2-7 | 184,619 |
33
+ | **Text-only capacity/default** | **`compose.sm120-tp2-language-only.yaml`** | **none; built-in MTP3** | **1,376,256** |
34
+
35
+ The MTP3 option means the model's built-in MTP head only: it does not load or
36
+ mount the external DFlash checkpoint. Choose DFlash2 when its modest C1 decode
37
+ gain matters more than resident context/concurrency; choose MTP3 for the normal
38
+ text-only daily driver.
39
 
40
  ## Run the current image
41
 
 
105
  ./serve-glm53-sm120-tp2-language-only.sh
106
  ```
107
 
108
+ To keep DFlash2 while disabling vision, use the separate speed-first launcher:
109
+
110
+ ```bash
111
+ curl -L -o compose.sm120-tp2-language-only-dflash2.yaml \
112
+ https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2-language-only-dflash2.yaml
113
+
114
+ GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
115
+ GLM53_DFLASH_PATH=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
116
+ docker compose -f compose.sm120-tp2-language-only-dflash2.yaml up -d
117
+ ```
118
+
119
+ Its standalone equivalent is
120
+ [`serve-glm53-sm120-tp2-language-only-dflash2.sh`](runtime/serve-glm53-sm120-tp2-language-only-dflash2.sh).
121
+
122
  The language-only alias points to the same tested v84 code digest:
123
 
124
  ```text
 
137
  Turning vision off raises the DFlash KV token pool by 42.6%. The much larger
138
  7.45x language-only gain comes from using the built-in MTP head instead of
139
  keeping the external DFlash2-7 model resident. It is not a vision-only gain.
140
+ The language-only profiles do not accept image inputs. The reported KV-token
141
+ pool is total allocated capacity, not a promise that every request can use the
142
+ entire pool; the configured per-request ceiling and scheduler concurrency still
143
+ apply.
144
 
145
  ## Current measured results
146
 
runtime/compose.sm120-tp2-language-only-dflash2.yaml ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ services:
2
+ glm53-flash-language-only-dflash2:
3
+ image: verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only@sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692
4
+ container_name: glm53-flash-exl3-k4-language-only-dflash2
5
+ init: true
6
+ ipc: host
7
+ shm_size: 32gb
8
+ restart: unless-stopped
9
+ ports:
10
+ - "${GLM53_PORT:-8012}:${GLM53_PORT:-8012}"
11
+ environment:
12
+ VLLM_ENGINE_READY_TIMEOUT_S: "3600"
13
+ VLLM_B12X_GLM_NOPE_NVFP4: "1"
14
+ VLLM_NVFP4_MLA_DYNAMIC_SCALE: "0"
15
+ VLLM_NVFP4_MLA_SCALES_FILE: /opt/glm53/calibration/glm53_nvfp4_mla_outer_scales_mtp_power2_v2.json
16
+ VLLM_EXL3_PREFILL_BLOCK_M: "128"
17
+ VLLM_EXL3_PREFILL_TRELLIS: "1"
18
+ B12X_GL53_ROUTE128_WIDE: "1"
19
+ B12X_GL53_ROUTE128_HYBRID_TAIL: "1"
20
+ VLLM_USE_B12X_DCP_A2A: "1"
21
+ VLLM_ENABLE_PCIE_ALLREDUCE: "1"
22
+ VLLM_PCIE_ALLREDUCE_BACKEND: cpp
23
+ KV_FP8_ROPE: "0"
24
+ OMP_NUM_THREADS: "2"
25
+ NCCL_IB_DISABLE: "1"
26
+ NCCL_P2P_LEVEL: "4"
27
+ volumes:
28
+ - "${GLM53_MODEL_PATH:?set GLM53_MODEL_PATH to the EXL3 checkpoint}:/model:ro"
29
+ - "${GLM53_DFLASH_PATH:?set GLM53_DFLASH_PATH to incoai/GLM-5.3-Flash-DFlash2}:/draft:ro"
30
+ - "${GLM53_CACHE_PATH:-./glm53-vllm-cache}:/cache"
31
+ command:
32
+ - serve
33
+ - /model
34
+ - --served-model-name
35
+ - GLM-5.3-Flash-EXL3-4bpw
36
+ - --host
37
+ - 0.0.0.0
38
+ - --port
39
+ - "${GLM53_PORT:-8012}"
40
+ - --language-model-only
41
+ - --tensor-parallel-size
42
+ - "2"
43
+ - --enable-expert-parallel
44
+ - --decode-context-parallel-size
45
+ - "2"
46
+ - --dcp-comm-backend
47
+ - a2a
48
+ - --dtype
49
+ - bfloat16
50
+ - --load-format
51
+ - safetensors
52
+ - --moe-backend
53
+ - b12x
54
+ - --attention-backend
55
+ - B12X_MLA_SPARSE
56
+ - --kv-cache-dtype
57
+ - nvfp4_ds_mla
58
+ - --max-model-len
59
+ - "98304"
60
+ - --max-num-batched-tokens
61
+ - "2072"
62
+ - --max-num-seqs
63
+ - "4"
64
+ - --gpu-memory-utilization
65
+ - "0.986"
66
+ - --enable-chunked-prefill
67
+ - --no-enable-prefix-caching
68
+ - --generation-config
69
+ - /model
70
+ - --reasoning-parser
71
+ - glm45
72
+ - --tool-call-parser
73
+ - glm47
74
+ - --enable-auto-tool-choice
75
+ - --disable-custom-all-reduce
76
+ - --speculative-config
77
+ - '{"method":"dflash","model":"/draft","num_speculative_tokens":7,"draft_tensor_parallel_size":2,"draft_sample_method":"probabilistic","rejection_sample_method":"standard","attention_backend":"TRITON_ATTN","kv_cache_dtype":"auto"}'
78
+ deploy:
79
+ resources:
80
+ reservations:
81
+ devices:
82
+ - driver: nvidia
83
+ device_ids: ["${GLM53_GPU_0:-0}", "${GLM53_GPU_1:-1}"]
84
+ capabilities: [gpu]
runtime/serve-glm53-sm120-tp2-language-only-dflash2.sh ADDED
@@ -0,0 +1,60 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env bash
2
+ set -euo pipefail
3
+
4
+ IMAGE="${IMAGE:-verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only@sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692}"
5
+ MODEL="${MODEL:?set MODEL to the local EXL3 checkpoint directory}"
6
+ DFLASH_MODEL="${DFLASH_MODEL:?set DFLASH_MODEL to the local incoai/GLM-5.3-Flash-DFlash2 directory}"
7
+ GPU_DEVICES="${GPU_DEVICES:-0,1}"
8
+ PORT="${PORT:-8012}"
9
+ NAME="${NAME:-glm53-flash-exl3-k4-language-only-dflash2}"
10
+ CACHE_PATH="${GLM53_CACHE_PATH:-${PWD}/glm53-vllm-cache}"
11
+
12
+ mkdir -p "${CACHE_PATH}"
13
+
14
+ exec docker run --rm --name "${NAME}" \
15
+ --init --gpus "\"device=${GPU_DEVICES}\"" --ipc=host --shm-size 32g \
16
+ -p "${PORT}:${PORT}" \
17
+ -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
18
+ -e VLLM_B12X_GLM_NOPE_NVFP4=1 \
19
+ -e VLLM_NVFP4_MLA_DYNAMIC_SCALE=0 \
20
+ -e VLLM_NVFP4_MLA_SCALES_FILE=/opt/glm53/calibration/glm53_nvfp4_mla_outer_scales_mtp_power2_v2.json \
21
+ -e VLLM_EXL3_PREFILL_BLOCK_M=128 \
22
+ -e VLLM_EXL3_PREFILL_TRELLIS=1 \
23
+ -e B12X_GL53_ROUTE128_WIDE=1 \
24
+ -e B12X_GL53_ROUTE128_HYBRID_TAIL=1 \
25
+ -e VLLM_USE_B12X_DCP_A2A=1 \
26
+ -e VLLM_ENABLE_PCIE_ALLREDUCE=1 \
27
+ -e VLLM_PCIE_ALLREDUCE_BACKEND=cpp \
28
+ -e KV_FP8_ROPE=0 \
29
+ -e OMP_NUM_THREADS=2 \
30
+ -e NCCL_IB_DISABLE=1 \
31
+ -e NCCL_P2P_LEVEL=4 \
32
+ -v "${MODEL}:/model:ro" \
33
+ -v "${DFLASH_MODEL}:/draft:ro" \
34
+ -v "${CACHE_PATH}:/cache" \
35
+ "${IMAGE}" serve /model \
36
+ --served-model-name GLM-5.3-Flash-EXL3-4bpw \
37
+ --host 0.0.0.0 --port "${PORT}" \
38
+ --language-model-only \
39
+ --tensor-parallel-size 2 \
40
+ --enable-expert-parallel \
41
+ --decode-context-parallel-size 2 \
42
+ --dcp-comm-backend a2a \
43
+ --dtype bfloat16 \
44
+ --load-format safetensors \
45
+ --moe-backend b12x \
46
+ --attention-backend B12X_MLA_SPARSE \
47
+ --kv-cache-dtype nvfp4_ds_mla \
48
+ --max-model-len 98304 \
49
+ --max-num-batched-tokens 2072 \
50
+ --max-num-seqs 4 \
51
+ --gpu-memory-utilization 0.986 \
52
+ --enable-chunked-prefill \
53
+ --no-enable-prefix-caching \
54
+ --generation-config /model \
55
+ --reasoning-parser glm45 \
56
+ --tool-call-parser glm47 \
57
+ --enable-auto-tool-choice \
58
+ --disable-custom-all-reduce \
59
+ --speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":7,"draft_tensor_parallel_size":2,"draft_sample_method":"probabilistic","rejection_sample_method":"standard","attention_backend":"TRITON_ATTN","kv_cache_dtype":"auto"}' \
60
+ "$@"