--- license: other library_name: transformers pipeline_tag: text-generation base_model: - zai-org/GLM-5.3-BF16 base_model_relation: quantized gated: manual language: - en - zh tags: - glm - glm-5.3 - nvfp4 - modelopt - sglang - dflash2 - blackwell - de-risked - enterprise --- ![BLACKFROST RESEARCH](ASSETS/BLACKFROST-RESEARCH-BANNER.png) # GLM-5.3-DERISKED-NVFP4-V2 Deployment-focused ModelOpt NVFP4 derivative of the full GLM-5.3 Mixture-of-Experts model for NVIDIA Blackwell. It retains the complete expert topology and native tokenizer, chat template, reasoning, tool-use, long-context, and MTP metadata. This is a public, manually gated customer release. Blackfrost's weight-level production process is proprietary and is not distributed in this repository. ## Specifications | | | |---|---| | Architecture | `GlmMoeDsaForCausalLM` | | Precision | ModelOpt routed-expert W4A4 NVFP4 mixed precision | | Layers | 78 main layers plus the native MTP layer | | Experts | 256 routed experts, top-8 active, plus shared expert | | Hidden size | 6,144 | | Context ceiling | 1,048,576 positions | | Source | `zai-org/GLM-5.3-BF16` | | Target hardware | NVIDIA Blackwell | ## DFlash2 deployment The included deployment kit enables SGLang DFLASH speculative decoding with the separate revision-pinned `incoai/GLM-5.3-DFlash2` companion. Its weights are not redistributed here. Review the companion's CC BY-NC-ND 4.0 terms before use. See [`DEPLOYMENT/README.md`](DEPLOYMENT/README.md) and [`DEPLOYMENT/LAUNCH.sh`](DEPLOYMENT/LAUNCH.sh). Qualify the model for your workload before production. Operators remain responsible for access control, monitoring, legal compliance, and treating all model output as untrusted. The model is provided as is, without warranty.