ThinkingCap-Qwen3.8-27B — W4A16 "fast" variant for HyperQwen
Single-GPU (24 GB) W4A16 build of ThinkingCap-Qwen3.8-27B (BottleCap AI's reasoning fine-tune of Qwen3.8-27B), repackaged for the HyperQwen vLLM fork and tuned for single-user serving with MTP self-speculative decoding.
Derived from wasifb/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16 (base model): the AutoRound GPTQ weights are repacked in the compressed-tensors layout, the language-model head and MTP module are GPTQ-calibrated INT4, and the draft head is sliced to a 40k-token draft vocabulary.
Quantization
| Component | Format |
|---|---|
| Decoder linears | GPTQ INT4, group 128, symmetric |
| Language-model head | GPTQ INT4, group 128 (calibrated) |
| Token embeddings | INT8, group 128 (INT4 alternative below) |
| MTP draft module | GPTQ INT4, group 128 (calibrated) |
| Draft head | INT4, 40k-token draft vocab (mtp_draft_vocab_ids.pt) |
| Norms | BF16 |
| KV cache | int8 |
Serve
This checkpoint is built for the HyperQwen vLLM fork — serve it with the serving repo's single-user launcher:
hf download dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast --local-dir ./model-fast
MODEL=path/to/model-fast PREFIX_CACHE=1 VISION=1 SPEC=mtp CTX=long REQ_METRICS=1 bash single-user/start_qwen.sh
GPU KV cache size: 201,275 tokens — maximum concurrency for 150,000 tokens per request: 1.34x.
Smallest weights — optional INT4 embedding (fast-emb4)
The INT8 embedding above is the fastest. If you want the smallest footprint instead, the companion repo dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4 ships just the model-00002 shard rebuilt with an INT4 (group 128) embedding table (~0.6 GB smaller) plus the matching config. From your model directory:
hf download dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4 --local-dir ./emb4
cp ./emb4/model-00002-of-00005.safetensors ./emb4/config.json ./emb4/quantization_config.json .
The model then serves with the INT4 embedding (everything else — body, head, MTP — is unchanged, and the freed ~0.6 GB grows the KV pool — see the emb4 repo README for its serving numbers). To switch back, keep a copy of the original model-00002-of-00005.safetensors + config.json (or re-download this repo).
Hardware
One 24 GB GPU (RTX 3090/4090-class); ~15-16 GB weights + the 201,275-token int8 KV pool of the profile above.
License
PolyForm Small Business License 1.0.0 — the same license as the base model bottlecapai/ThinkingCap-Qwen3.8-27B, including the additional personal-use permission stated in LICENSE. Commercial use by organizations that do not qualify as a "Small Business" under that license requires a separate commercial license from BottleCap AI. The upstream Qwen materials remain Apache-2.0 (see NOTICE).
- Downloads last month
- 393