README: list quantized vs BF16 layers
Browse files
README.md
CHANGED
|
@@ -8,3 +8,20 @@ base_model_relation: quantized
|
|
| 8 |
Quantized version of [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
|
| 9 |
|
| 10 |
If you come here with older cards like A100, A6000 or 3090, consider using [vLLM](https://github.com/wtdcode/vllm-backport) which has served billions of tokens for frontier models with older GPUs.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
Quantized version of [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
|
| 9 |
|
| 10 |
If you come here with older cards like A100, A6000 or 3090, consider using [vLLM](https://github.com/wtdcode/vllm-backport) which has served billions of tokens for frontier models with older GPUs.
|
| 11 |
+
|
| 12 |
+
## What is quantized
|
| 13 |
+
|
| 14 |
+
Weight-only INT4 (symmetric, group size 128) via AWQ, stored in the compressed-tensors `pack-quantized` format.
|
| 15 |
+
|
| 16 |
+
**Quantized (INT4 W4A16):**
|
| 17 |
+
- Routed MoE experts in all 48 decoder layers: `mlp.experts.{0..511}.{gate_proj, up_proj, down_proj}` (≈123B of the 180B parameters)
|
| 18 |
+
|
| 19 |
+
**Kept in BF16 (not quantized):**
|
| 20 |
+
- Token embedding (`embed_tokens`) and `lm_head`
|
| 21 |
+
- Gated DeltaNet linear attention (`linear_attn.*`)
|
| 22 |
+
- Qwen Sparse Attention (`self_attn.{q,k,v,o}_proj`) and its indexer (`self_attn.indexer.*`)
|
| 23 |
+
- Gated residual / hyper-connections (`*_hyper_connection.*`, `hyper_connection_mixer.*`)
|
| 24 |
+
- MoE router (`mlp.gate`), shared expert (`mlp.shared_expert.*`) and its gate (`shared_expert_gate`)
|
| 25 |
+
- Per-layer embedding (PLE) block including the 51B n-gram embedding table (`ple.*`)
|
| 26 |
+
- Vision encoder (`model.visual.*`)
|
| 27 |
+
- MTP layer (`mtp.*`, in `model_mtp.safetensors`)
|