Image-Text-to-Text
Transformers
Safetensors
glm5_next
glm
exl3
tr3
vllm
sm120
nvfp4
dflash2
multimodal
shapleymcg
conversational
Eval Results (legacy)
4-bit precision
Instructions to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="brandonmusic/GLM-5.3-Flash-tr3-4bpw") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw") model = AutoModelForMultimodalLM.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.3-Flash-tr3-4bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
- SGLang
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
Document selectable MTP3 and DFlash2 profiles
Browse filesAdds directly runnable multimodal DFlash2, language-only DFlash2, and language-only built-in MTP3 options with the measured KV-token capacity difference.
README.md
CHANGED
|
@@ -19,9 +19,23 @@ tags:
|
|
| 19 |
# GLM-5.3-Flash TR3 4bpw — current SM120 runtime
|
| 20 |
|
| 21 |
This is the uniform-K4 EXL3/TR3 routed-expert checkpoint for GLM-5.3-Flash.
|
| 22 |
-
The current
|
| 23 |
-
|
| 24 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
## Run the current image
|
| 27 |
|
|
@@ -91,6 +105,20 @@ GPU_DEVICES=0,1 \
|
|
| 91 |
./serve-glm53-sm120-tp2-language-only.sh
|
| 92 |
```
|
| 93 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 94 |
The language-only alias points to the same tested v84 code digest:
|
| 95 |
|
| 96 |
```text
|
|
@@ -109,7 +137,10 @@ Measured capacity on the same two 96 GB GPUs at 300 W each:
|
|
| 109 |
Turning vision off raises the DFlash KV token pool by 42.6%. The much larger
|
| 110 |
7.45x language-only gain comes from using the built-in MTP head instead of
|
| 111 |
keeping the external DFlash2-7 model resident. It is not a vision-only gain.
|
| 112 |
-
The language-only
|
|
|
|
|
|
|
|
|
|
| 113 |
|
| 114 |
## Current measured results
|
| 115 |
|
|
|
|
| 19 |
# GLM-5.3-Flash TR3 4bpw — current SM120 runtime
|
| 20 |
|
| 21 |
This is the uniform-K4 EXL3/TR3 routed-expert checkpoint for GLM-5.3-Flash.
|
| 22 |
+
The current v84 runtime supports three explicit TP2/EP2/DCP2 profiles on two
|
| 23 |
+
SM120 GPUs: multimodal DFlash2, language-only DFlash2, and language-only MTP3.
|
| 24 |
+
All use calibrated NVFP4 MLA KV and CUDA graphs. This is a custom vLLM/B12X
|
| 25 |
+
build and is not compatible with stock upstream vLLM.
|
| 26 |
+
|
| 27 |
+
## Pick a serving profile
|
| 28 |
+
|
| 29 |
+
| Goal | Launcher | Extra checkpoint | Measured KV tokens |
|
| 30 |
+
|---|---|---|---:|
|
| 31 |
+
| Images plus fastest measured C1 decode | `compose.sm120-tp2.yaml` | DFlash2-7 | 129,473 |
|
| 32 |
+
| Text-only DFlash2 decode | `compose.sm120-tp2-language-only-dflash2.yaml` | DFlash2-7 | 184,619 |
|
| 33 |
+
| **Text-only capacity/default** | **`compose.sm120-tp2-language-only.yaml`** | **none; built-in MTP3** | **1,376,256** |
|
| 34 |
+
|
| 35 |
+
The MTP3 option means the model's built-in MTP head only: it does not load or
|
| 36 |
+
mount the external DFlash checkpoint. Choose DFlash2 when its modest C1 decode
|
| 37 |
+
gain matters more than resident context/concurrency; choose MTP3 for the normal
|
| 38 |
+
text-only daily driver.
|
| 39 |
|
| 40 |
## Run the current image
|
| 41 |
|
|
|
|
| 105 |
./serve-glm53-sm120-tp2-language-only.sh
|
| 106 |
```
|
| 107 |
|
| 108 |
+
To keep DFlash2 while disabling vision, use the separate speed-first launcher:
|
| 109 |
+
|
| 110 |
+
```bash
|
| 111 |
+
curl -L -o compose.sm120-tp2-language-only-dflash2.yaml \
|
| 112 |
+
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2-language-only-dflash2.yaml
|
| 113 |
+
|
| 114 |
+
GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
|
| 115 |
+
GLM53_DFLASH_PATH=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
|
| 116 |
+
docker compose -f compose.sm120-tp2-language-only-dflash2.yaml up -d
|
| 117 |
+
```
|
| 118 |
+
|
| 119 |
+
Its standalone equivalent is
|
| 120 |
+
[`serve-glm53-sm120-tp2-language-only-dflash2.sh`](runtime/serve-glm53-sm120-tp2-language-only-dflash2.sh).
|
| 121 |
+
|
| 122 |
The language-only alias points to the same tested v84 code digest:
|
| 123 |
|
| 124 |
```text
|
|
|
|
| 137 |
Turning vision off raises the DFlash KV token pool by 42.6%. The much larger
|
| 138 |
7.45x language-only gain comes from using the built-in MTP head instead of
|
| 139 |
keeping the external DFlash2-7 model resident. It is not a vision-only gain.
|
| 140 |
+
The language-only profiles do not accept image inputs. The reported KV-token
|
| 141 |
+
pool is total allocated capacity, not a promise that every request can use the
|
| 142 |
+
entire pool; the configured per-request ceiling and scheduler concurrency still
|
| 143 |
+
apply.
|
| 144 |
|
| 145 |
## Current measured results
|
| 146 |
|
runtime/compose.sm120-tp2-language-only-dflash2.yaml
ADDED
|
@@ -0,0 +1,84 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
services:
|
| 2 |
+
glm53-flash-language-only-dflash2:
|
| 3 |
+
image: verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only@sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692
|
| 4 |
+
container_name: glm53-flash-exl3-k4-language-only-dflash2
|
| 5 |
+
init: true
|
| 6 |
+
ipc: host
|
| 7 |
+
shm_size: 32gb
|
| 8 |
+
restart: unless-stopped
|
| 9 |
+
ports:
|
| 10 |
+
- "${GLM53_PORT:-8012}:${GLM53_PORT:-8012}"
|
| 11 |
+
environment:
|
| 12 |
+
VLLM_ENGINE_READY_TIMEOUT_S: "3600"
|
| 13 |
+
VLLM_B12X_GLM_NOPE_NVFP4: "1"
|
| 14 |
+
VLLM_NVFP4_MLA_DYNAMIC_SCALE: "0"
|
| 15 |
+
VLLM_NVFP4_MLA_SCALES_FILE: /opt/glm53/calibration/glm53_nvfp4_mla_outer_scales_mtp_power2_v2.json
|
| 16 |
+
VLLM_EXL3_PREFILL_BLOCK_M: "128"
|
| 17 |
+
VLLM_EXL3_PREFILL_TRELLIS: "1"
|
| 18 |
+
B12X_GL53_ROUTE128_WIDE: "1"
|
| 19 |
+
B12X_GL53_ROUTE128_HYBRID_TAIL: "1"
|
| 20 |
+
VLLM_USE_B12X_DCP_A2A: "1"
|
| 21 |
+
VLLM_ENABLE_PCIE_ALLREDUCE: "1"
|
| 22 |
+
VLLM_PCIE_ALLREDUCE_BACKEND: cpp
|
| 23 |
+
KV_FP8_ROPE: "0"
|
| 24 |
+
OMP_NUM_THREADS: "2"
|
| 25 |
+
NCCL_IB_DISABLE: "1"
|
| 26 |
+
NCCL_P2P_LEVEL: "4"
|
| 27 |
+
volumes:
|
| 28 |
+
- "${GLM53_MODEL_PATH:?set GLM53_MODEL_PATH to the EXL3 checkpoint}:/model:ro"
|
| 29 |
+
- "${GLM53_DFLASH_PATH:?set GLM53_DFLASH_PATH to incoai/GLM-5.3-Flash-DFlash2}:/draft:ro"
|
| 30 |
+
- "${GLM53_CACHE_PATH:-./glm53-vllm-cache}:/cache"
|
| 31 |
+
command:
|
| 32 |
+
- serve
|
| 33 |
+
- /model
|
| 34 |
+
- --served-model-name
|
| 35 |
+
- GLM-5.3-Flash-EXL3-4bpw
|
| 36 |
+
- --host
|
| 37 |
+
- 0.0.0.0
|
| 38 |
+
- --port
|
| 39 |
+
- "${GLM53_PORT:-8012}"
|
| 40 |
+
- --language-model-only
|
| 41 |
+
- --tensor-parallel-size
|
| 42 |
+
- "2"
|
| 43 |
+
- --enable-expert-parallel
|
| 44 |
+
- --decode-context-parallel-size
|
| 45 |
+
- "2"
|
| 46 |
+
- --dcp-comm-backend
|
| 47 |
+
- a2a
|
| 48 |
+
- --dtype
|
| 49 |
+
- bfloat16
|
| 50 |
+
- --load-format
|
| 51 |
+
- safetensors
|
| 52 |
+
- --moe-backend
|
| 53 |
+
- b12x
|
| 54 |
+
- --attention-backend
|
| 55 |
+
- B12X_MLA_SPARSE
|
| 56 |
+
- --kv-cache-dtype
|
| 57 |
+
- nvfp4_ds_mla
|
| 58 |
+
- --max-model-len
|
| 59 |
+
- "98304"
|
| 60 |
+
- --max-num-batched-tokens
|
| 61 |
+
- "2072"
|
| 62 |
+
- --max-num-seqs
|
| 63 |
+
- "4"
|
| 64 |
+
- --gpu-memory-utilization
|
| 65 |
+
- "0.986"
|
| 66 |
+
- --enable-chunked-prefill
|
| 67 |
+
- --no-enable-prefix-caching
|
| 68 |
+
- --generation-config
|
| 69 |
+
- /model
|
| 70 |
+
- --reasoning-parser
|
| 71 |
+
- glm45
|
| 72 |
+
- --tool-call-parser
|
| 73 |
+
- glm47
|
| 74 |
+
- --enable-auto-tool-choice
|
| 75 |
+
- --disable-custom-all-reduce
|
| 76 |
+
- --speculative-config
|
| 77 |
+
- '{"method":"dflash","model":"/draft","num_speculative_tokens":7,"draft_tensor_parallel_size":2,"draft_sample_method":"probabilistic","rejection_sample_method":"standard","attention_backend":"TRITON_ATTN","kv_cache_dtype":"auto"}'
|
| 78 |
+
deploy:
|
| 79 |
+
resources:
|
| 80 |
+
reservations:
|
| 81 |
+
devices:
|
| 82 |
+
- driver: nvidia
|
| 83 |
+
device_ids: ["${GLM53_GPU_0:-0}", "${GLM53_GPU_1:-1}"]
|
| 84 |
+
capabilities: [gpu]
|
runtime/serve-glm53-sm120-tp2-language-only-dflash2.sh
ADDED
|
@@ -0,0 +1,60 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
set -euo pipefail
|
| 3 |
+
|
| 4 |
+
IMAGE="${IMAGE:-verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only@sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692}"
|
| 5 |
+
MODEL="${MODEL:?set MODEL to the local EXL3 checkpoint directory}"
|
| 6 |
+
DFLASH_MODEL="${DFLASH_MODEL:?set DFLASH_MODEL to the local incoai/GLM-5.3-Flash-DFlash2 directory}"
|
| 7 |
+
GPU_DEVICES="${GPU_DEVICES:-0,1}"
|
| 8 |
+
PORT="${PORT:-8012}"
|
| 9 |
+
NAME="${NAME:-glm53-flash-exl3-k4-language-only-dflash2}"
|
| 10 |
+
CACHE_PATH="${GLM53_CACHE_PATH:-${PWD}/glm53-vllm-cache}"
|
| 11 |
+
|
| 12 |
+
mkdir -p "${CACHE_PATH}"
|
| 13 |
+
|
| 14 |
+
exec docker run --rm --name "${NAME}" \
|
| 15 |
+
--init --gpus "\"device=${GPU_DEVICES}\"" --ipc=host --shm-size 32g \
|
| 16 |
+
-p "${PORT}:${PORT}" \
|
| 17 |
+
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
|
| 18 |
+
-e VLLM_B12X_GLM_NOPE_NVFP4=1 \
|
| 19 |
+
-e VLLM_NVFP4_MLA_DYNAMIC_SCALE=0 \
|
| 20 |
+
-e VLLM_NVFP4_MLA_SCALES_FILE=/opt/glm53/calibration/glm53_nvfp4_mla_outer_scales_mtp_power2_v2.json \
|
| 21 |
+
-e VLLM_EXL3_PREFILL_BLOCK_M=128 \
|
| 22 |
+
-e VLLM_EXL3_PREFILL_TRELLIS=1 \
|
| 23 |
+
-e B12X_GL53_ROUTE128_WIDE=1 \
|
| 24 |
+
-e B12X_GL53_ROUTE128_HYBRID_TAIL=1 \
|
| 25 |
+
-e VLLM_USE_B12X_DCP_A2A=1 \
|
| 26 |
+
-e VLLM_ENABLE_PCIE_ALLREDUCE=1 \
|
| 27 |
+
-e VLLM_PCIE_ALLREDUCE_BACKEND=cpp \
|
| 28 |
+
-e KV_FP8_ROPE=0 \
|
| 29 |
+
-e OMP_NUM_THREADS=2 \
|
| 30 |
+
-e NCCL_IB_DISABLE=1 \
|
| 31 |
+
-e NCCL_P2P_LEVEL=4 \
|
| 32 |
+
-v "${MODEL}:/model:ro" \
|
| 33 |
+
-v "${DFLASH_MODEL}:/draft:ro" \
|
| 34 |
+
-v "${CACHE_PATH}:/cache" \
|
| 35 |
+
"${IMAGE}" serve /model \
|
| 36 |
+
--served-model-name GLM-5.3-Flash-EXL3-4bpw \
|
| 37 |
+
--host 0.0.0.0 --port "${PORT}" \
|
| 38 |
+
--language-model-only \
|
| 39 |
+
--tensor-parallel-size 2 \
|
| 40 |
+
--enable-expert-parallel \
|
| 41 |
+
--decode-context-parallel-size 2 \
|
| 42 |
+
--dcp-comm-backend a2a \
|
| 43 |
+
--dtype bfloat16 \
|
| 44 |
+
--load-format safetensors \
|
| 45 |
+
--moe-backend b12x \
|
| 46 |
+
--attention-backend B12X_MLA_SPARSE \
|
| 47 |
+
--kv-cache-dtype nvfp4_ds_mla \
|
| 48 |
+
--max-model-len 98304 \
|
| 49 |
+
--max-num-batched-tokens 2072 \
|
| 50 |
+
--max-num-seqs 4 \
|
| 51 |
+
--gpu-memory-utilization 0.986 \
|
| 52 |
+
--enable-chunked-prefill \
|
| 53 |
+
--no-enable-prefix-caching \
|
| 54 |
+
--generation-config /model \
|
| 55 |
+
--reasoning-parser glm45 \
|
| 56 |
+
--tool-call-parser glm47 \
|
| 57 |
+
--enable-auto-tool-choice \
|
| 58 |
+
--disable-custom-all-reduce \
|
| 59 |
+
--speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":7,"draft_tensor_parallel_size":2,"draft_sample_method":"probabilistic","rejection_sample_method":"standard","attention_backend":"TRITON_ATTN","kv_cache_dtype":"auto"}' \
|
| 60 |
+
"$@"
|