How to use from
SGLang
# Gated model: Login with a HF token with gated access permission
hf auth login
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Quick Links

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

BLACKFROST RESEARCH

GLM-5.3-DERISKED-NVFP4-V2

Deployment-focused ModelOpt NVFP4 derivative of the full GLM-5.3 Mixture-of-Experts model for NVIDIA Blackwell. It retains the complete expert topology and native tokenizer, chat template, reasoning, tool-use, long-context, and MTP metadata.

This is a public, manually gated customer release. Blackfrost's weight-level production process is proprietary and is not distributed in this repository.

Specifications

Architecture GlmMoeDsaForCausalLM
Precision ModelOpt routed-expert W4A4 NVFP4 mixed precision
Layers 78 main layers plus the native MTP layer
Experts 256 routed experts, top-8 active, plus shared expert
Hidden size 6,144
Context ceiling 1,048,576 positions
Source zai-org/GLM-5.3-BF16
Target hardware NVIDIA Blackwell

DFlash2 deployment

The included deployment kit enables SGLang DFLASH speculative decoding with the separate revision-pinned incoai/GLM-5.3-DFlash2 companion. Its weights are not redistributed here. Review the companion's CC BY-NC-ND 4.0 terms before use.

See DEPLOYMENT/README.md and DEPLOYMENT/LAUNCH.sh.

Qualify the model for your workload before production. Operators remain responsible for access control, monitoring, legal compliance, and treating all model output as untrusted. The model is provided as is, without warranty.

Downloads last month
16
Safetensors
Model size
391B params
Tensor type
BF16
·
U8
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-Research/GLM-5.3-DERISKED-NVFP4-V2

Quantized
(28)
this model