Qwen3.8-27B-AWQ-W4A16-ASYM-HyperQwen
Qwen3.8-27B-AWQ-W4A16-ASYM prepared for
HyperQwen — Qwen3.8-27B served fast on one 24 GB card — the way the
TurboQwen image expects it. The body is untouched: int4 asymmetric
AWQ, group 128, zero points, quantized with llm-compressor from bf16 weights; the vision
tower, the SSM gate projections and the MTP head's norms stay bf16. What changed is what
HyperQwen's prepare/ pipeline changes, so that a 24 GB card has room for a KV cache and
speculative decoding has something small to score:
| tensor | in the source export | here |
|---|---|---|
lm_head |
bf16, 2.5 GB | int8 (group 128, symmetric), round-trip error 0.7% (prepare/quant_heads_stream.py) |
embed_tokens |
bf16, 2.5 GB | int8 (group 128, symmetric), round-trip error 0.65% |
MTP module (mtp.*, 8 linears incl. mtp.fc) |
bf16, 850 MB | int8 (group 128, symmetric) |
mtp.draft_lm_head + mtp_draft_vocab_ids.pt |
— | 40,960-row draft head (int8, sliced from the int8 lm_head), HyperQwen's shipped id list |
| body (64 layers), vision tower | int4 asym g128 / bf16 | unchanged |
Qwen3.8-27B-AWQ-W4A16-ASYM-HyperQwen-fast is the same with an int4-GPTQ lm_head: the single-user "fast" variant, a few tok/s more.
Serving
The container does everything (download, verify, serve), with the vision tower on:
git clone https://github.com/Ar4ikov/TurboQwen && cd TurboQwen
cp .env.example .env # CHECKPOINT=base
docker compose --profile single up -d
Bare metal, on the fork branch this was measured with (Ar4ikov/HyperQwen@awq-asym,
vLLM 0.29.0 + the series + marlin-int8-asym-zp):
hf download Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM-HyperQwen --local-dir models/Qwen3.8-27B-AWQ-W4A16-ASYM-HyperQwen
VISION=1 MODEL=$PWD/models/Qwen3.8-27B-AWQ-W4A16-ASYM-HyperQwen SPEC=mtp CTX=fast bash single-user/start_qwen.sh
Plain vLLM 0.29 also loads it (vllm serve Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM-HyperQwen --max-model-len 65536), without the
speculative decoding, the draft head or the int8 Marlin path that need the patch series.
Measured
RTX 3090 (350 W), vLLM 0.29.0, HyperQwen bench/run_benchmarks.sh single, second run
after boot, VISION=1 (the tower streamed from host RAM per image). C1 = one stream of
real prompts with 1,024-token answers; tok/step = tokens accepted per forward pass.
| profile | C1 T=default | C1 T=0 | tok/step | C8 T=default | KV pool |
|---|---|---|---|---|---|
SPEC=mtp CTX=fast (64k, bf16 KV) |
111.4 tok/s | 116.0 | 2.85 / 2.82 | 437 tok/s | 70,933 |
Images are described correctly in every profile (a drawn red square, blue circle and a
line of text: boost/image_smoke.py). The whole table, the int8 (W4A8) profiles and the
kernel measurements: github.com/Ar4ikov/TurboQwen.
Why a separate repo
HyperQwen's pipeline rewrites the checkpoint in place (int8 heads, the draft head) and its launchers, verify script and Docker entrypoint expect that layout. Doing it once and publishing the result turns a ~10-minute CPU step per machine into a download, and keeps the original export as it is for transformers and plain vLLM users.
Quantization recipe of the body, calibration set and the AWQ mappings for the hybrid
Gated-DeltaNet / attention layers: see the source export's card and its recipe.yaml
(kept here). Root model: Qwen/Qwen3.8-27B, Apache-2.0.
- Downloads last month
- 103