Capicua25x commited on
Commit
57c7280
·
verified ·
1 Parent(s): e971741

Note: '8-bit precision' badge is HF reading uint8 packing — model is 4-bit MXFP4

Browse files
Files changed (1) hide show
  1. README.md +5 -0
README.md CHANGED
@@ -116,6 +116,11 @@ modules to `quantization_config.ignore`** (else vLLM loads them as quantized →
116
 
117
  ## Caveats
118
 
 
 
 
 
 
119
  - Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
120
  - Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
121
  - `--language-model-only` required (see Serving).
 
116
 
117
  ## Caveats
118
 
119
+ - **HF shows an "8-bit precision" badge — ignore it, this model is 4-bit.** MXFP4 packs two
120
+ 4-bit FP4 values into each `uint8` byte (`weight_packed`), so HF reads the `uint8` *storage*
121
+ dtype and mislabels it. Source of truth: `config.json` → `num_bits: 4`,
122
+ `format: mxfp4-pack-quantized`. Every MXFP4 compressed-tensors model shows this (incl. the base).
123
+
124
  - Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
125
  - Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
126
  - `--language-model-only` required (see Serving).