GLM-5.3-Flash W4A16 NVFP4 GGUF

GGUF conversion of axiomofmind/GLM-5.3-Flash-W4A16-NVFP4, based on zai-org/GLM-5.3-Flash-BF16.

The main-model routed experts use W4A16 NVFP4 with group size 16. Attention, shared experts, routers, embeddings, the output head, MTP weights, and other retained tensors remain in BF16 or F32.

Files

File Description Size
GLM-5.3-Flash-W4A16-NVFP4-BF16attn-MTP.gguf Text model with MTP weights 204.1 GB

Requirements

A llama.cpp build with GLM5Next and NVFP4 GGUF support is required.

License

This model is distributed under the MIT License. Refer to the official model card for architecture details, usage guidance, and limitations.

Downloads last month
403
GGUF
Model size
321B params
Architecture
glm5next
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF

Quantized
(1)
this model

Collection including axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF