Qwen3.8-27B Quantizations (INT4 + FP8-KV)
INT4 quantizations of Qwen3.8-27B calibrated at 262k context: AWQ and GPTQ weights, each in bf16-KV and calibrated fp8-e4m3 KV editions.
Image-Text-to-Text • 28B • Updated • 5.17k • 2Note This repository documents the KV-scaling methodology and the fp8-vs-bf16 side-by-side evaluation; its weights are byte-identical to its bf16-KV twin, so the deltas isolate the KV dtype.
abhishekchohan/Qwen3.8-27B-AWQ-INT4
Image-Text-to-Text • 28B • Updated • 1.57kNote The fp8-e4m3 KV edition of this checkpoint is abhishekchohan/Qwen3.8-27B-AWQ-INT4-FP8KV: byte-identical weights plus calibrated per-layer KV scales, with side-by-side evals that isolate the KV dtype.
abhishekchohan/Qwen3.8-27B-GPTQ-INT4-FP8KV
Image-Text-to-Text • 28B • Updated • 62.5k • 2Note GPTQ's Hessian path is stable at 8 packed sequences (~2.1M tokens); the AWQ editions use a larger 112-sequence set (~29.4M) because activation-aware scaling is more data-hungry. For fp8-vs-bf16 KV side-by-side evals that isolate the KV dtype, see the repository abhishekchohan/AWQ-INT4-FP8KV.
abhishekchohan/Qwen3.8-27B-GPTQ-INT4
Image-Text-to-Text • 28B • Updated • 525Note This is the earlier config that protects all linear_attn.* modules (kept BF16). The other three editions in this collection quantize the DeltaNet projections. For fp8-vs-bf16 KV side-by-side evals that isolate the KV dtype, see the repository abhishekchohan/AWQ-INT4-FP8KV.