Update README.md
Browse files
README.md
CHANGED
|
@@ -21,9 +21,9 @@ Samples were drawn from a diverse mix of publicly available datasets spanning co
|
|
| 21 |
|
| 22 |
### How to Run
|
| 23 |
|
| 24 |
-
|
| 25 |
|
| 26 |
-
|
| 27 |
|
| 28 |
Tested on 2x and 4x RTX Pro 6000 Blackwell.
|
| 29 |
|
|
@@ -55,4 +55,59 @@ python3 -m sglang.launch_server \
|
|
| 55 |
--port 8000
|
| 56 |
|
| 57 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
```
|
|
|
|
| 21 |
|
| 22 |
### How to Run
|
| 23 |
|
| 24 |
+
If you experience NCCL hangs with P2P, make sure you have `iommu=pt` (and `amd_iommu=pt` on AMD platforms) in your kernel command line.
|
| 25 |
|
| 26 |
+
#### SGLang
|
| 27 |
|
| 28 |
Tested on 2x and 4x RTX Pro 6000 Blackwell.
|
| 29 |
|
|
|
|
| 55 |
--port 8000
|
| 56 |
|
| 57 |
|
| 58 |
+
```
|
| 59 |
+
|
| 60 |
+
#### vLLM
|
| 61 |
+
|
| 62 |
+
(thanks to @zenmagnets)
|
| 63 |
+
|
| 64 |
+
Set your Hugging Face cache and GPUs, then run (from project root with venv activated).
|
| 65 |
+
```
|
| 66 |
+
export CUDA_DEVICE_ORDER=PCI_BUS_ID
|
| 67 |
+
export CUDA_VISIBLE_DEVICES=0,1
|
| 68 |
+
export HF_HOME=/path/to/huggingface
|
| 69 |
+
export HUGGINGFACE_HUB_CACHE=$HF_HOME/hub
|
| 70 |
+
export VLLM_WORKER_MULTIPROC_METHOD=spawn
|
| 71 |
+
export SAFETENSORS_FAST_GPU=1
|
| 72 |
+
export VLLM_NVFP4_GEMM_BACKEND=cutlass
|
| 73 |
+
export VLLM_USE_FLASHINFER_MOE_FP4=0
|
| 74 |
+
export VLLM_DISABLE_PYNCCL=1
|
| 75 |
+
export NCCL_IB_DISABLE=1
|
| 76 |
+
export OMP_NUM_THREADS=8
|
| 77 |
+
export VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
|
| 78 |
+
|
| 79 |
+
python -m vllm.entrypoints.openai.api_server \
|
| 80 |
+
--model lukealonso/MiniMax-M2.5-NVFP4 \
|
| 81 |
+
--download-dir $HUGGINGFACE_HUB_CACHE \
|
| 82 |
+
--host 0.0.0.0 \
|
| 83 |
+
--port 1235 \
|
| 84 |
+
--served-model-name MiniMax-M2.5-NVFP4 \
|
| 85 |
+
--trust-remote-code \
|
| 86 |
+
--tensor-parallel-size 2 \
|
| 87 |
+
--attention-backend FLASH_ATTN \
|
| 88 |
+
--gpu-memory-utilization 0.95 \
|
| 89 |
+
--max-model-len 190000 \
|
| 90 |
+
--max-num-batched-tokens 16384 \
|
| 91 |
+
--max-num-seqs 64 \
|
| 92 |
+
--disable-custom-all-reduce \
|
| 93 |
+
--enable-auto-tool-choice \
|
| 94 |
+
--tool-call-parser minimax_m2 \
|
| 95 |
+
--reasoning-parser minimax_m2_append_think
|
| 96 |
+
```
|
| 97 |
+
|
| 98 |
+
Dependencies
|
| 99 |
+
|
| 100 |
+
Install in a Python 3.12 venv; use CUDA 12.x on the host.
|
| 101 |
+
|
| 102 |
+
```
|
| 103 |
+
Package Version Note
|
| 104 |
+
vllm 0.15.1 OpenAI server + NVFP4 MoE
|
| 105 |
+
torch 2.9.1+cu128 CUDA 12.8 build
|
| 106 |
+
transformers 4.57.6
|
| 107 |
+
safetensors 0.7.0
|
| 108 |
+
nvidia-modelopt 0.41.0 NVFP4 / ModelOpt format
|
| 109 |
+
flashinfer-python 0.6.1 Optional (we use FLASH_ATTN)
|
| 110 |
+
nvidia-nccl-cu12 2.27.5 Multi-GPU
|
| 111 |
+
nvidia-cutlass-dsl* 4.4.0.dev1 NVFP4 GEMM (script uses cutlass backend)
|
| 112 |
+
System: CUDA 12.8, cuDNN 9.10.2 (or matching torch cuDNN). Driver must support your GPUs (e.g. Blackwell).
|
| 113 |
```
|