--- base_model: "MiniMaxAI/MiniMax-M2.5" --- ## Model Description **MiniMax-M2.5-NVFP4** is an NVFP4-quantized version of [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5), a 456B-parameter Mixture-of-Experts language model with 46B active parameters. The original model weights were converted from the official FP8 checkpoint to BF16, then quantized to NVFP4 (4-bit with blockwise FP8 scales per 16 elements) using [NVIDIA TensorRT Model Optimizer](https://github.com/NVIDIA/TensorRT-Model-Optimizer). ### What's quantized Only the MoE expert MLP layers (gate, up, and down projections) are quantized to NVFP4. Attention layers are left in BF16. Since the expert weights constitute the vast majority of model parameters in an MoE architecture, this still yields significant memory savings. Calibration uses natural top-k routing rather than forcing all experts to activate, so each expert's quantization scales reflect the token distributions it actually sees during inference. To compensate, calibration was run on a significantly larger number of samples than typical to ensure broad expert coverage through natural routing alone. ### Calibration dataset Samples were drawn from a diverse mix of publicly available datasets spanning code generation, function/tool calling, multi-turn reasoning, math, and multilingual (English + Chinese) instruction following. System prompts were randomly varied across samples. The dataset was designed to broadly exercise the model's capabilities and activate diverse token distributions across expert modules. ### How to Run I've had bad luck recently with NVFP4 and vLLM, so I've personally switched to sglang. If someone has a working vLLM recipe, please share, but the below works fine for me on the latest sglang. ``` export NCCL_IB_DISABLE=1 export NCCL_P2P_LEVEL=PHB export NCCL_ALLOC_P2P_NET_LL_BUFFERS=1 export NCCL_MIN_NCHANNELS=8 export OMP_NUM_THREADS=8 export SAFETENSORS_FAST_GPU=1 python3 -m sglang.launch_server \ --model lukealonso/MiniMax-M2.5-NVFP4 \ --trust-remote-code \ --tp 8 \ --mem-fraction-static 0.9 \ --max-running-requests 16 \ --attention-backend flashinfer \ --kv-cache-dtype fp8_e4m3 \ --moe-runner-backend flashinfer_cutlass \ --disable-custom-all-reduce \ --enable-flashinfer-allreduce-fusion \ --host 0.0.0.0 \ --port 8000 ```