Download README.md from lukealonso/MiniMax-M2.5-NVFP4: direct link, hf CLI and curl.
- Browser
- Download file 2.49 kB
-
https://huggingface.co/lukealonso/MiniMax-M2.5-NVFP4/resolve/8dede464b5c64bfece40ce4c9077f099c1107a79/README.md
- Command line
-
hf download hf://lukealonso/MiniMax-M2.5-NVFP4@8dede464b5c64bfece40ce4c9077f099c1107a79/README.md
-
curl -L -o README.md https://huggingface.co/lukealonso/MiniMax-M2.5-NVFP4/resolve/8dede464b5c64bfece40ce4c9077f099c1107a79/README.md
base_model:
- MiniMaxAI/MiniMax-M2.5
Model Description
MiniMax-M2.5-NVFP4 is an NVFP4-quantized version of MiniMaxAI/MiniMax-M2.5, a 456B-parameter Mixture-of-Experts language model with 46B active parameters.
The original model weights were converted from the official FP8 checkpoint to BF16, then quantized to NVFP4 (4-bit with blockwise FP8 scales per 16 elements) using NVIDIA Model Optimizer.
What's quantized
Only the MoE expert MLP layers (gate, up, and down projections) are quantized to NVFP4. Attention layers are left in BF16. Since the expert weights constitute the vast majority of model parameters in an MoE architecture, this still yields significant memory savings.
Calibration uses natural top-k routing rather than forcing all experts to activate, so each expert's quantization scales reflect the token distributions it actually sees during inference. To compensate, calibration was run on a significantly larger number of samples than typical to ensure broad expert coverage through natural routing alone.
Calibration dataset
Samples were drawn from a diverse mix of publicly available datasets spanning code generation, function/tool calling, multi-turn reasoning, math, and multilingual (English + Chinese) instruction following. System prompts were randomly varied across samples. The dataset was designed to broadly exercise the model's capabilities and activate diverse token distributions across expert modules.
How to Run
I've had bad luck recently with NVFP4 and vLLM, so I've personally switched to sglang.
If someone has a working vLLM recipe, please share, but the below works fine for me on the latest sglang.
export NCCL_IB_DISABLE=1
export NCCL_P2P_LEVEL=PHB
export NCCL_ALLOC_P2P_NET_LL_BUFFERS=1
export NCCL_MIN_NCHANNELS=8
export OMP_NUM_THREADS=8
export SAFETENSORS_FAST_GPU=1
python3 -m sglang.launch_server \
--model /data/models/MiniMax-M2.5-NVFP4 \
--served-model-name MiniMax-M2.5-NVFP4 \
--reasoning-parser minimax \
--enable-torch-compile \
--trust-remote-code \
--tp 2--ep 2 \
--mem-fraction-static 0.9 \
--max-running-requests 16 \
--kv-cache-dtype fp8_e4m3 \
--quantization modelopt_fp4 \
--attention-backend flashinfer \
--moe-runner-backend flashinfer_cutlass \
--disable-custom-all-reduce \
--enable-flashinfer-allreduce-fusion \
--host 0.0.0.0 \
--port 8000