leoncca's picture
Clarify validation protocols and loading requirements
d9e14e5 verified
|
Raw
History Blame Contribute Delete
6.6 kB
---
base_model:
- Qwen/Qwen3.8-Flash-Next
base_model_relation: quantized
license: other
license_name: qwen-community-1.0
license_link: LICENSE
pipeline_tag: image-text-to-text
tags:
- qwen3.8
- qwen4-exp
- awq
- w4a16
- fp8
- e4m3
---
# Qwen3.8 Flash Next AWQ g32 + QSA FP8 E4M3 KV
AWQ derivative of
[`Qwen/Qwen3.8-Flash-Next`](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
Only routed-expert projections are quantized. PLE is reused from the official
FP8 checkpoint, and calibrated QSA FP8 E4M3 K/V scales are included.
**138.13 GB across 44 model shards: 10 AWQ, 33 PLE, and one K/V-scale shard.**
## Checkpoint
- Base revision: `f5d08274bafd880402bd16f5e3e6c514136ec06c`
- PLE revision: `bcd9f01ddc9cff2316eb84281bebcd5b058bddce`
- Routed experts: asymmetric AWQ W4A16, g32, zero point, GEMM layout
- Quantized: per-expert `gate_proj`, `up_proj`, and `down_proj`
- Kept at source precision: vision, PLE, attention, routers, shared experts,
embeddings, LM head, hyper-connections, and MTP
- Indexed tensors: 222,771 with 137,042,968,666 bytes of tensor payload
- Merged index SHA-256:
`f066a0154a9101b359c3a4d4fa6a83fb12b7126a9cf0f19cb772611d08cc07ee`
Machine-readable details are in `EXPERT_COVERAGE.json`,
`MODEL_PROVENANCE.json`, `VALIDATION.json`, and `SHA256SUMS`.
## Why AWQ g32
Two properties favor asymmetric AWQ here:
1. On hardware without native FP4 execution, AWQ W4A16 provides a practical
weight-only path while keeping activations in FP16.
2. The routed-expert weights frequently have off-center local ranges, which an
asymmetric affine codebook can represent with a movable zero point.
A route-frequency-stratified CPU audit covered all 93 layer/expert pairs with
fewer than 128 natural tokens, including all eight zero-hit pairs, plus controls
from every layer. In total it compared 707 layer/expert pairs, all three routed
projections, 2,121 matrices, and 69,500,928 BF16 weight values. It dequantized
the actual NVFP4 checkpoint and compared both formats against the same BF16
mother weights; all 48 unquantized router tensors matched bitwise.
| Route stratum | Pairs | Matrices | Actual NVFP4 relative RMSE | Affine g32 relative RMSE | Reduction | g32 wins |
|---|---:|---:|---:|---:|---:|---:|
| Zero natural tokens | 8 | 24 | 0.094704 | 0.082320 | 13.08% | 24/24 |
| 1–31 natural tokens | 14 | 42 | 0.094387 | 0.083253 | 11.80% | 42/42 |
| 32–63 natural tokens | 17 | 51 | 0.094737 | 0.082029 | 13.41% | 51/51 |
| 64–127 natural tokens | 54 | 162 | 0.094747 | 0.082104 | 13.34% | 162/162 |
| Complete `<128` census | 93 | 279 | 0.094685 | 0.082289 | 13.09% | 279/279 |
| Non-low-frequency controls | 614 | 1,842 | 0.094887 | 0.081588 | 14.02% | 1,842/1,842 |
| All deduplicated selections | 707 | 2,121 | 0.094859 | 0.081686 | 13.89% | 2,121/2,121 |
Affine g32 had lower error in every sampled matrix, including all 279
low-frequency gate/up/down matrices. The complete low-frequency census also
reduced NMSE by 24.47%.
The result was not only a group-size effect: asymmetric affine g16 reached
0.069948 relative RMSE, compared with 0.094859 for actual NVFP4. Of the sampled
BF16 groups, 56.78% had a range-center shift above 0.1 and 25.00% above 0.2.
Only 55.40% of raw affine g32 zero-points and 58.73% of released AWQ zero-points
landed at the central 7/8 positions. These weights benefit from a movable
asymmetric zero point rather than a fixed zero-centered codebook.
The g32 layout also matches the model's width-640 experts under TP4:
`640 / 4 = 160`, and 160 is divisible by 32 but not 128. This avoids storing
g128 metadata that must later be expanded to effective g32. The cost is about
6.59 GiB more scale and zero-point metadata than g128.
## Calibration coverage
AWQ weight calibration used 684 records and 202,750 active tokens, producing
97,320,000 native top-10 token-to-expert assignments across 48 layers. Natural
routing covered 24,568 of 24,576 layer/expert pairs (`99.9674%`), with 45/48
layers at 512/512. The remaining eight zero-hit pairs and 85 low-coverage pairs
were explicitly augmented for gate/up/down projection calibration. The released
checkpoint requires zero runtime fallback pairs. This corpus is separate from
the 4,130,597-token E4M3 K/V-scale calibration described below.
## Loading requirements
The loader must support:
- Qwen4Exp per-expert asymmetric AWQ W4A16 g32 in GEMM layout;
- indexed reuse of complete PLE shard files while ignoring unindexed tensors;
- all 24 bundled, finite, positive QSA K/V scale entries;
- complete Qwen4Exp multimodal support, including MRoPE and video-token-pruning
initialization, for image and video inference.
The `image-text-to-text` tag requires the complete multimodal path; text-only
loading does not establish image or video support.
## Validation
| Quality check | FP16 KV | Calibrated E4M3 KV |
|---|---:|---:|
| Basic generation | 4/4 | — |
| Needle retrieval, length sweep | 6/6: one case each at 1K, 4K, 16K, 32K, 64K, and 128K | — |
| Needle retrieval, 128K repeat set | — | 6/6: three cases × two repeats |
| Held-out tool selection | 10/12 | 10/12 |
| General repeat set | 6/6 exact-stable | — |
| Held-out tool repeat set | 3/6 exact-stable | 4/6 exact-stable |
| GSM8K five-shot subset | 29/32 | 29/32 |
| HumanEval/MBPP functional subset | 9/10 | 9/10 |
| IFEval | 3/5 prompts, 9/12 instructions | 3/5 prompts, 9/12 instructions |
| Image and video requests | 18/18 semantic; 9/9 cases exact-stable | 18/18 semantic; 9/9 cases exact-stable |
`—` means that the check was not run in that K/V arm; it is not a failed result.
The E4M3 scales were calibrated from 4,130,597 tokens. All 24 values are finite
and range from `0.0171072837` to `0.0806361660`, with zero observed calibration
saturation. First tokens matched on 18/18 checks. Selected-block recall averaged
`0.995906` with a minimum of `0.991822`; QSA outputs had minimum cosine
`0.998597` and maximum relative L2 `0.052962`.
The multimodal set contained five image and four video cases, each repeated
twice. All 333 `model.visual.*` tensors are bitwise identical to the reference
checkpoint, so AWQ did not alter the vision tower payload.
## Notes
- The bundled K/V scales belong to this exact merged checkpoint and should not
be replaced with scales from another weight artifact.
- The reported results are regression checks rather than a full capability
benchmark.
- This derivative uses the included Qwen Community License 1.0.
## Integrity
Verify the repository with `SHA256SUMS`. The authoritative artifact identity is
the merged index SHA-256 shown above.