ThinkingCap-Qwen3.8-27B W4A16 — INT4 embedding shard (companion to "-fast")
Partial companion repo for
dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast:
just the model-00002 shard rebuilt with an INT4 (group 128) embedding table instead of
INT8 (~0.6 GB smaller on disk/VRAM), plus the matching config.json /
quantization_config.json (identical otherwise — same INT4 GPTQ body, head and MTP as the
main repo).
This is not a standalone model — it exists to switch the main model's embedding
quantization. From your main model directory (after hf download-ing the -fast repo):
hf download dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4 --local-dir ./emb4
cp ./emb4/model-00002-of-00005.safetensors ./emb4/config.json ./emb4/quantization_config.json .
# restart the server
To switch back to the INT8 embedding: restore the original
model-00002-of-00005.safetensors + config.json from the main repo (keep a copy before
overwriting, or re-download it).
Serve
Serve the swapped model (e.g. a directory named model-fast-emb4) with the
HyperQwen vLLM fork single-user launcher:
MODEL=path/to/model-fast-emb4 PREFIX_CACHE=1 VISION=1 SPEC=mtp CTX=long REQ_METRICS=1 bash single-user/start_qwen.sh
GPU KV cache size: 218,112 tokens — maximum concurrency for 150,000 tokens per request: 1.45x (the INT4 embedding frees ~0.6 GB back to the KV pool vs the INT8 version).
License
PolyForm Small Business License 1.0.0, same as the base model and the main repo — see LICENSE and NOTICE (and the main repo's README for provenance).
- Downloads last month
- 335
Model tree for dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4
Base model
Qwen/Qwen3.8-27B