bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ

NVFP4 (W4A4) quantization of bottlecapai/ThinkingCap-Qwen3.8-27B, a Qwen3.8-27B fine-tune, built with llm-compressor and stored in the compressed-tensors format vLLM loads natively. 23.4 GB on disk, against ≈56 GB for the bf16 source.

The MTP (multi-token-prediction) head and the vision tower are kept in bf16 and work: self-speculative decoding and image input both serve from this repo.

What is quantized

Weights were smoothed with AWQ (activation-aware quantization) before rounding. The counts below are the language model's Linear layers; the vision tower and the MTP head are bf16 throughout and are not part of them.

precision linear layers which
4-bit (NVFP4) 168 the MLP projections (gate_proj / up_proj / down_proj) of all but the last 8 layers
FP8 (per-tensor, W8A8) 233 self-attention (q/k/v/o_proj), the wide Gated-DeltaNet projections (in_proj_qkv / in_proj_z / out_proj), the last 8 layers' MLPs, and lm_head
bf16 96 the two narrow Gated-DeltaNet input gates in each linear-attention layer (linear_attn.in_proj_a / in_proj_b, 48-wide — not a multiple of the 4-bit kernel's tile size, and fused to width 96 at serve time, so quantizing them breaks loading)

A mixed 4-bit/FP8 layout rather than 4-bit everywhere: the attention and linear-attention projections are where a flat 4-bit quantization of this architecture loses accuracy, and FP8 costs little memory over 4-bit at that share of the weights.

Hardware

Needs Blackwell (sm_100 / sm_120: RTX PRO 6000, B200, GB200). Weights and activations are FP4, which is what reaches the FP4 tensor cores — vLLM serves this through CutlassNvFp4LinearKernel. This build's numbers on this card come from an RTX PRO 6000; on other GPUs serve NVFP4 (4-bit weights, 16-bit activations) instead.

Serving

Tested on vLLM 0.29.0.

vllm serve bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --max-model-len 49152 \
  --max-num-seqs 32

The model's native context is 262,144 tokens; --max-model-len above is a memory-frugal default, not a limit of the checkpoint.

--max-num-seqs matters here: this is a hybrid Gated-DeltaNet model, which allocates one state cache block per decode sequence, and vLLM's default of 1024 overruns them and aborts CUDA-graph capture.

Speculative decoding (MTP)

The model's own NextN head drafts for it — no draft model to download:

vllm serve bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --max-model-len 49152 --max-num-seqs 32 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}'

Pin the drafter's attention backend. Left unset, vLLM picks FlashInfer for the MTP head, whose reorder-batch threshold (1) is taken as the minimum against the hybrid target's Triton backend (4); decode rows then queue behind continuing prefill chunks, the speculative state slots are never initialised, and long generations degenerate into repeated function words partway through — silently, without an error (vLLM #55894).

Sampling

Thinking mode, as the base model: temperature 1.0, top_p 0.95, top_k 20, min_p 0.0.

Expected performance

Paired comparison with the bf16 source on the full quantization plan: both builds answer the same questions with the same seeds — RealWorldQA 765 questions × 2 seeds (images), GPQA-Diamond 198 × 4, MMLU-Pro 1,500 × 1 (a fixed slice), IFBench 300 × 2, AA-LCR 100 × 1 (89k–123k-token prompts, graded by Gemma-4-26B-A4B-it with thinking off). Thinking at the chat template's default reasoning effort (xhigh), sampled decoding (temperature 1.0, top_p 0.95, top_k 20, min_p 0.0), 65,536-token generation cap, vLLM 0.29.0. This build's answers were generated on an RTX PRO 6000 Blackwell, the bf16 reference on an H200.

benchmark (questions × seeds) accuracy %, bf16 → NVFP4A4-AWQ Δ accuracy, pp [95% CI] tokens mean / median / p95, bf16 → NVFP4A4-AWQ Δ mean tokens [95% CI] Δ median tokens
RealWorldQA (765 × 2) 83.1 → 81.8 −1.3 [−2.9, +0.3] 488 / 112 / 1,912 → 506 / 114 / 2,093 +3.7% [−11.2, +20.6] +1.8%
GPQA-Diamond (198 × 4) 88.0 → 86.7 −1.3 [−3.6, +1.1] 7,115 / 1,031 / 37,459 → 7,176 / 1,156 / 38,429 +0.9% [−6.8, +9.2] +12.1%
MMLU-Pro (1,500 × 1) 84.1 → 84.3 +0.2 [−1.1, +1.5] 1,436 / 166 / 7,255 → 1,487 / 177 / 7,068 +3.5% [−9.7, +19.4] +6.6%
IFBench (300 × 2) 79.7 → 79.7 0.0 [−2.9, +2.9] 4,531 / 1,822 / 20,630 → 4,387 / 1,718 / 17,549 −3.2% [−11.6, +6.9] −5.7%
AA-LCR (100 × 1) 81.0 → 82.0 +1.0 [−6.1, +8.1] 1,718 / 844 / 4,937 → 1,798 / 1,148 / 5,284 +4.6% [−8.1, +19.2] +36.0%

Δ accuracy is this build minus bf16 on the same answers; its interval treats the question as the unit (seeds averaged per question first). Tokens are completion tokens (reasoning plus answer).

GPQA-Diamond seed 1 was generated with a 131,072-token cap and truncated to 65,536 afterwards; 2 of its 198 answers end at the cap.

Throughput — one RTX PRO 6000 Blackwell, vLLM 0.29.0, synthetic prompts of 1,024 tokens with 512 generated (last column: 32,768-token prompts, 128 generated), end-of-sequence ignored, prefix caching off, --max-num-seqs 64. Aggregate output tokens/s, median time to first token in ms in parentheses.

build 1 request 16 concurrent 64 concurrent 4 concurrent, 32k-token prompts
bf16 26.2 (158) 341 (1,867) 844 (3,627) 18.6 (9,678)
NVFP4 (weight-only, Marlin), served once 68.0 (176) 666 (2,404) 1,149 (4,105) 18.9 (13,670)
NVFP4A4-AWQ (CUTLASS FP4) 48.1 (82) 623 (935) 1,435 (1,773) 32.8 (6,393)

Against NVFP4, the weight-only 4-bit build, it is 25% faster at 64 concurrent requests, 74% faster with a 53% lower median time to first token on 32k-token prompts, and 29% slower for a single request. Pick it for batched serving on Blackwell; for one request at a time, or on a GPU without FP4 tensor cores, serve NVFP4.

Where to find us

Website LinkedIn Instagram X

Need even more efficiency? The open release is production-ready. Our enterprise versions go further — fewer thinking tokens still, tuned to your workload, at matched accuracy on your own tasks. Built for AI labs, inference providers and enterprises running models at scale. Deployed on your infrastructure, or in the cloud and region you choose. Talk to our team

License

ThinkingCap: PolyForm Small Business 1.0.0 + BottleCap personal-use grant (see LICENSE).

Upstream Qwen materials: Apache-2.0 (see NOTICE).

Commercial license: contact BottleCap AI.

Citation

If you use this model, please cite:

@misc{ThinkingCap-Qwen3.8-27B,
  title     = {bottlecapai/ThinkingCap-Qwen3.8-27B},
  author    = {Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and Bartek, Vojtech and Jirak, Jiri and Kubista, Daniel and Krus, Frantisek and Mikolov, Tomas},
  year      = {2026},
}
Downloads last month
472
Safetensors
Model size
20B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ

Base model

Qwen/Qwen3.8-27B
Quantized
(19)
this model

Collection including bottlecapai/ThinkingCap-Qwen3.8-27B-NVFP4A4-AWQ