Image-Text-to-Text
Transformers
Safetensors
glm5_next
glm
exl3
tr3
vllm
sm120
nvfp4
dflash2
multimodal
shapleymcg
conversational
Eval Results (legacy)
4-bit precision
Instructions to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="brandonmusic/GLM-5.3-Flash-tr3-4bpw") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw") model = AutoModelForMultimodalLM.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.3-Flash-tr3-4bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
- SGLang
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
Link complete BF16 teacher logits dataset
Browse files
README.md
CHANGED
|
@@ -20,6 +20,24 @@ The direct packed TP2 serving result (`0.022750847878`) is a separate one-window
|
|
| 20 |
|
| 21 |
Code and the five-run receipts: [brandonmmusic-max/glm-5.3-flash-exl3-4bpw](https://github.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw).
|
| 22 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
## Minimal TP2 launch
|
| 24 |
Note this is NOT an optimize launch nearly at all. I woudl recommend trying https://github.com/chriswritescode-dev/glm-5.3-flash-sm120 image referenced
|
| 25 |
here. that use the MLA which compresses the latent before you get to kv_b_proj. Getting that running you might get 1 million kv cache.
|
|
|
|
| 20 |
|
| 21 |
Code and the five-run receipts: [brandonmmusic-max/glm-5.3-flash-exl3-4bpw](https://github.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw).
|
| 22 |
|
| 23 |
+
## BF16 teacher logits and replay calibration
|
| 24 |
+
|
| 25 |
+
The complete teacher dataset is published at
|
| 26 |
+
[`brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits`](https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits).
|
| 27 |
+
It contains 640 rolling calibration windows plus the 25 qualification-only final
|
| 28 |
+
windows: 665 windows total, each with 2,048 input tokens, 2,047 scored positions,
|
| 29 |
+
and the full 154,880-token vocabulary. The 1,361,255 scored positions occupy
|
| 30 |
+
843,324,965,136 raw logits bytes.
|
| 31 |
+
|
| 32 |
+
The teacher is the released BF16 checkpoint with its native FP32 tensors
|
| 33 |
+
preserved; the logits are stored as float32 to avoid an additional
|
| 34 |
+
storage-precision loss. Payload revision
|
| 35 |
+
[`7c378d5f17dba158c4c803eff27c346dd0615660`](https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits/tree/7c378d5f17dba158c4c803eff27c346dd0615660)
|
| 36 |
+
is bound by the
|
| 37 |
+
[`16e16e90078bc0b54bd1cd37b08ba7dad03819726d0258443a0e30b68b354472` aggregate audit](https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits/blob/267ccf27ca92575529e0a1ef80e7eed8d209a8f4/logits/full-panel/receipts/full-dataset-audit.json),
|
| 38 |
+
which records every payload path, size, and SHA-256. The 25 final windows remain
|
| 39 |
+
qualification-only and are excluded from fitting and expert selection.
|
| 40 |
+
|
| 41 |
## Minimal TP2 launch
|
| 42 |
Note this is NOT an optimize launch nearly at all. I woudl recommend trying https://github.com/chriswritescode-dev/glm-5.3-flash-sm120 image referenced
|
| 43 |
here. that use the MLA which compresses the latent before you get to kv_b_proj. Getting that running you might get 1 million kv cache.
|