Instructions to use Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2") model = AutoModelForCausalLM.from_pretrained("Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2
- SGLang
How to use Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2 with Docker Model Runner:
docker model run hf.co/Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2
GLM-5.3-DERISKED-NVFP4-V2
Deployment-focused ModelOpt NVFP4 derivative of the full GLM-5.3 Mixture-of-Experts model for NVIDIA Blackwell. It retains the complete expert topology and native tokenizer, chat template, reasoning, tool-use, long-context, and MTP metadata.
This is a public, manually gated customer release. Blackfrost's weight-level production process is proprietary and is not distributed in this repository.
Specifications
| Architecture | GlmMoeDsaForCausalLM |
| Precision | ModelOpt routed-expert W4A4 NVFP4 mixed precision |
| Layers | 78 main layers plus the native MTP layer |
| Experts | 256 routed experts, top-8 active, plus shared expert |
| Hidden size | 6,144 |
| Context ceiling | 1,048,576 positions |
| Source | zai-org/GLM-5.3-BF16 |
| Target hardware | NVIDIA Blackwell |
DFlash2 deployment
The included deployment kit enables SGLang DFLASH speculative decoding with the
separate revision-pinned incoai/GLM-5.3-DFlash2 companion. Its weights are not
redistributed here. Review the companion's CC BY-NC-ND 4.0 terms before use.
See DEPLOYMENT/README.md and
DEPLOYMENT/LAUNCH.sh.
Qualify the model for your workload before production. Operators remain responsible for access control, monitoring, legal compliance, and treating all model output as untrusted. The model is provided as is, without warranty.
- Downloads last month
- 10
Model tree for Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2
Base model
zai-org/GLM-5.3-BF16