Note: '8-bit precision' badge is HF reading uint8 packing — model is 4-bit MXFP4
Browse files
README.md
CHANGED
|
@@ -116,6 +116,11 @@ modules to `quantization_config.ignore`** (else vLLM loads them as quantized →
|
|
| 116 |
|
| 117 |
## Caveats
|
| 118 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 119 |
- Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
|
| 120 |
- Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
|
| 121 |
- `--language-model-only` required (see Serving).
|
|
|
|
| 116 |
|
| 117 |
## Caveats
|
| 118 |
|
| 119 |
+
- **HF shows an "8-bit precision" badge — ignore it, this model is 4-bit.** MXFP4 packs two
|
| 120 |
+
4-bit FP4 values into each `uint8` byte (`weight_packed`), so HF reads the `uint8` *storage*
|
| 121 |
+
dtype and mislabels it. Source of truth: `config.json` → `num_bits: 4`,
|
| 122 |
+
`format: mxfp4-pack-quantized`. Every MXFP4 compressed-tensors model shows this (incl. the base).
|
| 123 |
+
|
| 124 |
- Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
|
| 125 |
- Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
|
| 126 |
- `--language-model-only` required (see Serving).
|