Update README.md
Browse files
README.md
CHANGED
|
@@ -13,7 +13,7 @@ The original model weights were converted from the official FP8 checkpoint to BF
|
|
| 13 |
|
| 14 |
Only the MoE expert MLP layers (gate, up, and down projections) are quantized to NVFP4. Attention layers are left in BF16. Since the expert weights constitute the vast majority of model parameters in an MoE architecture, this still yields significant memory savings.
|
| 15 |
|
| 16 |
-
Calibration uses natural top-k routing rather than forcing all experts to activate, so each expert's quantization scales reflect the token distributions it actually sees during inference. To compensate, calibration was run on a
|
| 17 |
|
| 18 |
### Calibration dataset
|
| 19 |
|
|
|
|
| 13 |
|
| 14 |
Only the MoE expert MLP layers (gate, up, and down projections) are quantized to NVFP4. Attention layers are left in BF16. Since the expert weights constitute the vast majority of model parameters in an MoE architecture, this still yields significant memory savings.
|
| 15 |
|
| 16 |
+
Calibration uses natural top-k routing rather than forcing all experts to activate, so each expert's quantization scales reflect the token distributions it actually sees during inference. To compensate, calibration was run on a vastly larger number of samples than typical to ensure broad expert coverage through natural routing alone.
|
| 17 |
|
| 18 |
### Calibration dataset
|
| 19 |
|