Update README.md
Browse files
README.md
CHANGED
|
@@ -3,8 +3,6 @@ base_model:
|
|
| 3 |
- MiniMaxAI/MiniMax-M2.5
|
| 4 |
---
|
| 5 |
|
| 6 |
-
(weights still uploading)
|
| 7 |
-
|
| 8 |
## Model Description
|
| 9 |
|
| 10 |
**MiniMax-M2.5-NVFP4** is an NVFP4-quantized version of [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5), a 456B-parameter Mixture-of-Experts language model with 46B active parameters.
|
|
@@ -28,6 +26,7 @@ I've had bad luck recently with NVFP4 and vLLM, so I've personally switched to s
|
|
| 28 |
If someone has a working vLLM recipe, please share, but the below works fine for me on the latest sglang.
|
| 29 |
|
| 30 |
```
|
|
|
|
| 31 |
export NCCL_IB_DISABLE=1
|
| 32 |
export NCCL_P2P_LEVEL=PHB
|
| 33 |
export NCCL_ALLOC_P2P_NET_LL_BUFFERS=1
|
|
@@ -36,17 +35,22 @@ export OMP_NUM_THREADS=8
|
|
| 36 |
export SAFETENSORS_FAST_GPU=1
|
| 37 |
|
| 38 |
python3 -m sglang.launch_server \
|
| 39 |
-
--model
|
|
|
|
|
|
|
|
|
|
| 40 |
--trust-remote-code \
|
| 41 |
-
--tp
|
| 42 |
--mem-fraction-static 0.9 \
|
| 43 |
--max-running-requests 16 \
|
| 44 |
-
--attention-backend flashinfer \
|
| 45 |
--kv-cache-dtype fp8_e4m3 \
|
|
|
|
|
|
|
| 46 |
--moe-runner-backend flashinfer_cutlass \
|
| 47 |
--disable-custom-all-reduce \
|
| 48 |
--enable-flashinfer-allreduce-fusion \
|
| 49 |
--host 0.0.0.0 \
|
| 50 |
--port 8000
|
| 51 |
|
|
|
|
| 52 |
```
|
|
|
|
| 3 |
- MiniMaxAI/MiniMax-M2.5
|
| 4 |
---
|
| 5 |
|
|
|
|
|
|
|
| 6 |
## Model Description
|
| 7 |
|
| 8 |
**MiniMax-M2.5-NVFP4** is an NVFP4-quantized version of [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5), a 456B-parameter Mixture-of-Experts language model with 46B active parameters.
|
|
|
|
| 26 |
If someone has a working vLLM recipe, please share, but the below works fine for me on the latest sglang.
|
| 27 |
|
| 28 |
```
|
| 29 |
+
|
| 30 |
export NCCL_IB_DISABLE=1
|
| 31 |
export NCCL_P2P_LEVEL=PHB
|
| 32 |
export NCCL_ALLOC_P2P_NET_LL_BUFFERS=1
|
|
|
|
| 35 |
export SAFETENSORS_FAST_GPU=1
|
| 36 |
|
| 37 |
python3 -m sglang.launch_server \
|
| 38 |
+
--model /data/models/MiniMax-M2.5-NVFP4 \
|
| 39 |
+
--served-model-name MiniMax-M2.5-NVFP4 \
|
| 40 |
+
--reasoning-parser minimax \
|
| 41 |
+
--enable-torch-compile \
|
| 42 |
--trust-remote-code \
|
| 43 |
+
--tp 2--ep 2 \
|
| 44 |
--mem-fraction-static 0.9 \
|
| 45 |
--max-running-requests 16 \
|
|
|
|
| 46 |
--kv-cache-dtype fp8_e4m3 \
|
| 47 |
+
--quantization modelopt_fp4 \
|
| 48 |
+
--attention-backend flashinfer \
|
| 49 |
--moe-runner-backend flashinfer_cutlass \
|
| 50 |
--disable-custom-all-reduce \
|
| 51 |
--enable-flashinfer-allreduce-fusion \
|
| 52 |
--host 0.0.0.0 \
|
| 53 |
--port 8000
|
| 54 |
|
| 55 |
+
|
| 56 |
```
|