lukealonso commited on
Commit
1f414f8
·
verified ·
1 Parent(s): 084c94d

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +57 -2
README.md CHANGED
@@ -21,9 +21,9 @@ Samples were drawn from a diverse mix of publicly available datasets spanning co
21
 
22
  ### How to Run
23
 
24
- I've had bad luck recently with NVFP4 and vLLM, so I've personally switched to sglang.
25
 
26
- If someone has a working vLLM recipe, please share, but the below works fine for me on the latest sglang.
27
 
28
  Tested on 2x and 4x RTX Pro 6000 Blackwell.
29
 
@@ -55,4 +55,59 @@ python3 -m sglang.launch_server \
55
  --port 8000
56
 
57
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
  ```
 
21
 
22
  ### How to Run
23
 
24
+ If you experience NCCL hangs with P2P, make sure you have `iommu=pt` (and `amd_iommu=pt` on AMD platforms) in your kernel command line.
25
 
26
+ #### SGLang
27
 
28
  Tested on 2x and 4x RTX Pro 6000 Blackwell.
29
 
 
55
  --port 8000
56
 
57
 
58
+ ```
59
+
60
+ #### vLLM
61
+
62
+ (thanks to @zenmagnets)
63
+
64
+ Set your Hugging Face cache and GPUs, then run (from project root with venv activated).
65
+ ```
66
+ export CUDA_DEVICE_ORDER=PCI_BUS_ID
67
+ export CUDA_VISIBLE_DEVICES=0,1
68
+ export HF_HOME=/path/to/huggingface
69
+ export HUGGINGFACE_HUB_CACHE=$HF_HOME/hub
70
+ export VLLM_WORKER_MULTIPROC_METHOD=spawn
71
+ export SAFETENSORS_FAST_GPU=1
72
+ export VLLM_NVFP4_GEMM_BACKEND=cutlass
73
+ export VLLM_USE_FLASHINFER_MOE_FP4=0
74
+ export VLLM_DISABLE_PYNCCL=1
75
+ export NCCL_IB_DISABLE=1
76
+ export OMP_NUM_THREADS=8
77
+ export VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
78
+
79
+ python -m vllm.entrypoints.openai.api_server \
80
+ --model lukealonso/MiniMax-M2.5-NVFP4 \
81
+ --download-dir $HUGGINGFACE_HUB_CACHE \
82
+ --host 0.0.0.0 \
83
+ --port 1235 \
84
+ --served-model-name MiniMax-M2.5-NVFP4 \
85
+ --trust-remote-code \
86
+ --tensor-parallel-size 2 \
87
+ --attention-backend FLASH_ATTN \
88
+ --gpu-memory-utilization 0.95 \
89
+ --max-model-len 190000 \
90
+ --max-num-batched-tokens 16384 \
91
+ --max-num-seqs 64 \
92
+ --disable-custom-all-reduce \
93
+ --enable-auto-tool-choice \
94
+ --tool-call-parser minimax_m2 \
95
+ --reasoning-parser minimax_m2_append_think
96
+ ```
97
+
98
+ Dependencies
99
+
100
+ Install in a Python 3.12 venv; use CUDA 12.x on the host.
101
+
102
+ ```
103
+ Package Version Note
104
+ vllm 0.15.1 OpenAI server + NVFP4 MoE
105
+ torch 2.9.1+cu128 CUDA 12.8 build
106
+ transformers 4.57.6
107
+ safetensors 0.7.0
108
+ nvidia-modelopt 0.41.0 NVFP4 / ModelOpt format
109
+ flashinfer-python 0.6.1 Optional (we use FLASH_ATTN)
110
+ nvidia-nccl-cu12 2.27.5 Multi-GPU
111
+ nvidia-cutlass-dsl* 4.4.0.dev1 NVFP4 GEMM (script uses cutlass backend)
112
+ System: CUDA 12.8, cuDNN 9.10.2 (or matching torch cuDNN). Driver must support your GPUs (e.g. Blackwell).
113
  ```