djdeniro commited on
Commit
1a7f400
·
verified ·
1 Parent(s): 7bfe6c8

Fix MTP info: not available in quantized variant; update performance

Browse files
Files changed (1) hide show
  1. README.md +2 -3
README.md CHANGED
@@ -39,7 +39,7 @@ The result fits in ~17.5 GiB per GPU (TP8) while retaining near-BF16 quality.
39
 
40
  - **456B total params** (sparse), **~30B activated** per token
41
  - **256 experts** per MoE layer, top-8 routing, 62 transformer layers
42
- - **3 MTP layers** for speculative decoding
43
  - **200k context window**
44
  - Native tool-calling support
45
 
@@ -81,10 +81,9 @@ docker run --name minimax-mxfp416 \
81
  -e ROCR_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
82
  -e TRUST_REMOTE_CODE=1 \
83
  -v /path/to/models:/app/models:ro \
84
- -v /path/to/patches:/patches:ro \
85
  -p 8000:8000 \
86
  tcclaviger/vllm22:latest \
87
- bash -c "cp /patches/vllm22_minimax_m2.py /app/vllm/vllm/model_executor/models/minimax_m2.py && \
88
  pip install -q sentencepiece && \
89
  exec vllm serve /app/models/MiniMax-M2.7-MXFP416 \
90
  --served-model-name minimax-m2.7-mxfp416 \
 
39
 
40
  - **456B total params** (sparse), **~30B activated** per token
41
  - **256 experts** per MoE layer, top-8 routing, 62 transformer layers
42
+ - Config declares `use_mtp: True` (1 layer, 3 modules), but MTP weights were stripped during quantization — not available in this variant
43
  - **200k context window**
44
  - Native tool-calling support
45
 
 
81
  -e ROCR_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
82
  -e TRUST_REMOTE_CODE=1 \
83
  -v /path/to/models:/app/models:ro \
 
84
  -p 8000:8000 \
85
  tcclaviger/vllm22:latest \
86
+ bash -c "cp /app/models/vllm22_minimax_m2.py /app/vllm/vllm/model_executor/models/minimax_m2.py && \
87
  pip install -q sentencepiece && \
88
  exec vllm serve /app/models/MiniMax-M2.7-MXFP416 \
89
  --served-model-name minimax-m2.7-mxfp416 \