hoborific commited on
Commit
747fd56
·
verified ·
1 Parent(s): a15b720

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +29 -0
README.md CHANGED
@@ -11,4 +11,33 @@ tags:
11
 
12
  Quantized version of [zerofata/G4-MeroMero-v2-31B](https://huggingface.co/zerofata/G4-MeroMero-v2-31B).
13
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
 
11
 
12
  Quantized version of [zerofata/G4-MeroMero-v2-31B](https://huggingface.co/zerofata/G4-MeroMero-v2-31B).
13
 
14
+ ## Format
15
+
16
+ Offline-quantized **W8A16 FP8** in the
17
+ [compressed-tensors](https://github.com/neuralmagic/compressed-tensors)
18
+ `float-quantized` format: weights in `float8_e4m3fn` with per-output-channel
19
+ symmetric scales, activations kept in bf16/fp16.
20
+
21
+ ## How it was quantized
22
+
23
+ For each linear layer, every output row gets its own scale starting from
24
+ `amax / 448`, refined by an MSE clip search over ~9 clip fractions
25
+ (0.8–1.0× amax) picking the lowest-error scale per row. Weights are then
26
+ quantized `q = e4m3(w / scale)` with round-to-nearest and saturation. This
27
+ per-channel + clipping scheme gives better SNR than vLLM's online per-tensor
28
+ `--quantization fp8` path.
29
+
30
+ Only 2D linear projection weights are quantized (attention q/k/v/o, MLP
31
+ gate/up/down). Embeddings, norms, lm_head, routers/experts, and the vision
32
+ tower stay in bf16 and are listed in the checkpoint's `ignore` list, so vLLM
33
+ leaves them untouched.
34
+
35
+ ## Supported vLLM platforms
36
+
37
+ - **Intel XPU** — `XPUW8A16FP8LinearKernel` (the intended target).
38
+ - **NVIDIA CUDA** (SM75+, i.e. Turing and newer) —
39
+ `HummingFP8ScaledMMLinearKernel` when the `humming` package is installed,
40
+ otherwise `MarlinFP8ScaledMMLinearKernel`.
41
+ - **Not supported**: ROCm, CPU, TPU — vLLM has no W8A16-FP8 kernel for these
42
+ backends yet, so loading will fail with a "no kernel" error.
43