wtdcode commited on
Commit
c7d2df0
·
verified ·
1 Parent(s): f60b002

README: list quantized vs BF16 layers

Browse files
Files changed (1) hide show
  1. README.md +17 -0
README.md CHANGED
@@ -8,3 +8,20 @@ base_model_relation: quantized
8
  Quantized version of [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
9
 
10
  If you come here with older cards like A100, A6000 or 3090, consider using [vLLM](https://github.com/wtdcode/vllm-backport) which has served billions of tokens for frontier models with older GPUs.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  Quantized version of [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
9
 
10
  If you come here with older cards like A100, A6000 or 3090, consider using [vLLM](https://github.com/wtdcode/vllm-backport) which has served billions of tokens for frontier models with older GPUs.
11
+
12
+ ## What is quantized
13
+
14
+ Weight-only INT4 (symmetric, group size 128) via AWQ, stored in the compressed-tensors `pack-quantized` format.
15
+
16
+ **Quantized (INT4 W4A16):**
17
+ - Routed MoE experts in all 48 decoder layers: `mlp.experts.{0..511}.{gate_proj, up_proj, down_proj}` (≈123B of the 180B parameters)
18
+
19
+ **Kept in BF16 (not quantized):**
20
+ - Token embedding (`embed_tokens`) and `lm_head`
21
+ - Gated DeltaNet linear attention (`linear_attn.*`)
22
+ - Qwen Sparse Attention (`self_attn.{q,k,v,o}_proj`) and its indexer (`self_attn.indexer.*`)
23
+ - Gated residual / hyper-connections (`*_hyper_connection.*`, `hyper_connection_mixer.*`)
24
+ - MoE router (`mlp.gate`), shared expert (`mlp.shared_expert.*`) and its gate (`shared_expert_gate`)
25
+ - Per-layer embedding (PLE) block including the 51B n-gram embedding table (`ple.*`)
26
+ - Vision encoder (`model.visual.*`)
27
+ - MTP layer (`mtp.*`, in `model_mtp.safetensors`)