ThinkingCap-Qwen3.8-27B W4A16 — INT4 embedding shard (companion to "-fast")

Partial companion repo for dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast: just the model-00002 shard rebuilt with an INT4 (group 128) embedding table instead of INT8 (~0.6 GB smaller on disk/VRAM), plus the matching config.json / quantization_config.json (identical otherwise — same INT4 GPTQ body, head and MTP as the main repo).

This is not a standalone model — it exists to switch the main model's embedding quantization. From your main model directory (after hf download-ing the -fast repo):

hf download dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4 --local-dir ./emb4
cp ./emb4/model-00002-of-00005.safetensors ./emb4/config.json ./emb4/quantization_config.json .
# restart the server

To switch back to the INT8 embedding: restore the original model-00002-of-00005.safetensors + config.json from the main repo (keep a copy before overwriting, or re-download it).

Serve

Serve the swapped model (e.g. a directory named model-fast-emb4) with the HyperQwen vLLM fork single-user launcher:

MODEL=path/to/model-fast-emb4 PREFIX_CACHE=1 VISION=1 SPEC=mtp CTX=long REQ_METRICS=1 bash single-user/start_qwen.sh

GPU KV cache size: 218,112 tokens — maximum concurrency for 150,000 tokens per request: 1.45x (the INT4 embedding frees ~0.6 GB back to the KV pool vs the INT8 version).

License

PolyForm Small Business License 1.0.0, same as the base model and the main repo — see LICENSE and NOTICE (and the main repo's README for provenance).

Downloads last month
335
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dewamade/ThinkingCap-Qwen3.8-27B-AutoRound-W4A16-fast-emb4