You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

ThinkingCap-Qwen3.8-27B — W4A16 "fast" variant for HyperQwen

Single-GPU (24 GB) W4A16 build of ThinkingCap-Qwen3.8-27B (BottleCap AI's reasoning fine-tune of Qwen3.8-27B), repackaged for the HyperQwen vLLM fork and tuned for single-user serving with MTP self-speculative decoding.

Derived from wasifb/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16 (base model): the AutoRound GPTQ weights are repacked in the compressed-tensors layout, the language-model head and MTP module are GPTQ-calibrated INT4, and the draft head is sliced to a 40k-token draft vocabulary.

Quantization

Component Format
Decoder linears GPTQ INT4, group 128, symmetric
Language-model head GPTQ INT4, group 128 (calibrated)
Token embeddings INT8, group 128 (INT4 alternative below)
MTP draft module GPTQ INT4, group 128 (calibrated)
Draft head INT4, 40k-token draft vocab (mtp_draft_vocab_ids.pt)
Norms BF16
KV cache int8

Serve

This checkpoint is built for the HyperQwen vLLM fork — serve it with the serving repo's single-user launcher:

hf download dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast --local-dir ./model-fast

MODEL=path/to/model-fast PREFIX_CACHE=1 VISION=1 SPEC=mtp CTX=long REQ_METRICS=1 bash single-user/start_qwen.sh

GPU KV cache size: 201,275 tokens — maximum concurrency for 150,000 tokens per request: 1.34x.

Smallest weights — optional INT4 embedding (fast-emb4)

The INT8 embedding above is the fastest. If you want the smallest footprint instead, the companion repo dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4 ships just the model-00002 shard rebuilt with an INT4 (group 128) embedding table (~0.6 GB smaller) plus the matching config. From your model directory:

hf download dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4 --local-dir ./emb4
cp ./emb4/model-00002-of-00005.safetensors ./emb4/config.json ./emb4/quantization_config.json .

The model then serves with the INT4 embedding (everything else — body, head, MTP — is unchanged, and the freed ~0.6 GB grows the KV pool — see the emb4 repo README for its serving numbers). To switch back, keep a copy of the original model-00002-of-00005.safetensors + config.json (or re-download this repo).

Hardware

One 24 GB GPU (RTX 3090/4090-class); ~15-16 GB weights + the 201,275-token int8 KV pool of the profile above.

License

PolyForm Small Business License 1.0.0 — the same license as the base model bottlecapai/ThinkingCap-Qwen3.8-27B, including the additional personal-use permission stated in LICENSE. Commercial use by organizations that do not qualify as a "Small Business" under that license requires a separate commercial license from BottleCap AI. The upstream Qwen materials remain Apache-2.0 (see NOTICE).

Downloads last month
393
Safetensors
Model size
3B params
Tensor type
BF16
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast

Base model

Qwen/Qwen3.8-27B
Quantized
(1)
this model
Quantizations
1 model