--- license: other library_name: transformers pipeline_tag: text-generation base_model: - zai-org/GLM-5.3-BF16 base_model_relation: quantized gated: manual language: - en - zh tags: - glm - glm-5.3 - fp8 - sglang - dflash2 - de-risked - enterprise --- ![BLACKFROST RESEARCH](ASSETS/BLACKFROST-RESEARCH-BANNER.png) # GLM-5.3-Derisked-FP8 Deployment-ready FP8 derivative of the full GLM-5.3 Mixture-of-Experts model. The model retains the complete expert topology and native GLM tokenizer, chat template, reasoning, tool-use, long-context, and MTP metadata. This is a public, manually gated customer release. Weight access is reviewed by Blackfrost Research and distributed under a separate commercial agreement. ## Specifications | | | |---|---| | Architecture | `GlmMoeDsaForCausalLM` | | Precision | Dynamic block-128 E4M3 FP8 | | Layers | 78 main layers plus the native MTP layer | | Experts | 256 routed experts, top-8 active, plus shared expert | | Hidden size | 6,144 | | Context ceiling | 1,048,576 positions | | Source | `zai-org/GLM-5.3-BF16` | Blackfrost's weight-level production process is proprietary and is not included in this repository. No adapter or client-side system prompt is required for the checkpoint's intended behavior. ## DFlash2 deployment The included deployment kit enables SGLang DFLASH speculative decoding with the separate revision-pinned `incoai/GLM-5.3-DFlash2` companion. Its weights are not redistributed here. Review the companion's CC BY-NC-ND 4.0 terms before use. See [`DEPLOYMENT/README.md`](DEPLOYMENT/README.md) and [`DEPLOYMENT/LAUNCH.sh`](DEPLOYMENT/LAUNCH.sh). ## Use and limitations Qualify capability, safety, long-context behavior, tool use, and speculative acceptance for your workload before production. Operators remain responsible for access control, monitoring, applicable-law compliance, and treating model output as untrusted. The model is provided as is, without warranty.