Download README.md from Shockem/Qwen3.8-27b-Terse-Coder-NVFP4: direct link, hf CLI and curl.
- Browser
- Download file 6.67 kB
-
https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder-NVFP4/resolve/4742480974c1b9d9fdaa03f7a5830d6e49f5fa9f/README.md
- Command line
-
hf download hf://Shockem/Qwen3.8-27b-Terse-Coder-NVFP4@4742480974c1b9d9fdaa03f7a5830d6e49f5fa9f/README.md
-
curl -L -o README.md https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder-NVFP4/resolve/4742480974c1b9d9fdaa03f7a5830d6e49f5fa9f/README.md
license: apache-2.0
base_model: Shockem/Qwen3.8-27b-Terse-Coder
base_model_relation: quantized
tags:
- reasoning
- coding
- qwen3
- nvfp4
- modelopt
Qwen3.8-27B Terse-Coder β NVFP4
NVFP4 (modelopt W4A16) quantization of Shockem/Qwen3.8-27b-Terse-Coder, a fine-tune of Qwen/Qwen3.8-27B with ~1/10 the chain-of-thought reasoning tokens on coding tasks and correctness preserved. This is the tested deployment artifact β every number below was measured on this checkpoint.
Actively researched and improving. Expect updated quants on this page as the study continues.
Results
Held-out 40 coding problems (20 HumanEval + 20 MBPP-sanitized, disjoint from training), vLLM 0.28 on 2Γ RTX 5060 Ti 16 GB, MTP spec decode on, sampling temp 0.6 / top_k 20 / top_p 0.95 / rep-penalty 1.05, pass@1 by automated test execution:
| Model (all NVFP4) | pass@1 | Reasoning tokens / problem | Wall tok/s |
|---|---|---|---|
| nvidia/Qwen3.8-27B-NVFP4 (stock) | 72.5% | ~701 | 54.1 |
| This model | 67.5% | ~38 (β95%) | 54.5 |
Runs at stock-base wall speed with MTP acceptance 0.412 β the reasoning cut
is free end-to-end. Independent benchmarks (NVFP4 quant, vLLM 0.28, thinking on, house
sampling; reasoning = completion_tokens_details.reasoning_tokens):
| Benchmark | Score | Reasoning tokens (mean / median) |
|---|---|---|
| GSM8K (n=200) | 98.0% | 84 / 72 |
| GPQA-Diamond (full 198) | 78.3% | 1,485 / 969 |
| CRUXEval-I (full 800, input prediction) | 92.1% | 197 / 83 |
| CRUXEval-O (full 800, output prediction) | 92.9% | 146 / 96 |
| HumanEval+ (164, official EvalPlus, greedy) | 90.2% (93.9% base) | 43 / 28 |
| MBPP+ (378, official EvalPlus, greedy) | 78.6% (92.9% base) | 91 / 25 |
CRUXEval was run with the official Meta harness (direct prompts, official extraction, exec-based scoring, temp 0.2) β code understanding (input/output prediction), complementing the generation-side coding table above.
A note on GPQA-Diamond: this is where a terseness fine-tune is supposed to bleed β PhD-level science, far outside the coding training distribution, where long deliberation is the whole game. Holding 78.3% at ~1.5k mean reasoning tokens (thinking models typically burn 10β20k here) means the training cut the deliberation budget, not the capability β the model still scales effort up on hard problems (median 969 β max 16k) instead of answering blindly fast.
Internal agentic harness (30 tests across easy/medium/hard β instruction following, coding, reasoning, compaction handoff, tool/JSON contracts β Γ10 runs each, this checkpoint served by vLLM): easy 100% (40/40), medium 100% (90/90), hard 100% (140/140), zero truncations, zero reasoning fallbacks. Prior best on the same harness was 100/100/98.7.
Quantization recipe
This is a v3-recipe house quant, built to preserve the adapter effect through 4-bit compression:
- modelopt 0.45 W4A16 NVFP4, per-tensor streaming PTQ (the same 400-tensor quantize set + ignore list as the published house Signal quants)
- FP8 attention (absmax β byte-matches NVIDIA's checkpoint at 97β99%)
- Local-Hessian-weighted calibration on MLP + lm_head (Hessian captured from 2048 house-traffic chunks; Hessian-weighted MSE scale solve with per-block e4m3 bracketing). This matters: an absmax-calibrated quant of the same weights attenuates the terse-reasoning effect to roughly half (β49.5% vs β92.4% cut measured). Geomean Hessian-weighted error ratio 0.805 vs the absmax baseline on the stock base.
- MTP draft stack included (1 MTP layer, BF16, vocab-truncated 40960-id draft head) so speculative decoding works out of the box.
Serving (vLLM, tested path)
vllm serve Shockem/Qwen3.8-27b-Terse-Coder-NVFP4 \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--kv-cache-dtype fp8
Turn MTP spec decode on β outputs are target-verified (lossless) and
acceptance is 0.41. If you serve with spec decode, make sure the generation
config has no min_p β vLLM 0.28 rejects min_p under spec decode.
Recommended sampling (mirrors testing): temp 0.6, top_k 20, top_p 0.95, repetition_penalty 1.05.
On 2Γ16 GB cards cap context at ~200k with a ~3.9 GiB FP8 KV pin; single-card 24 GB+ rigs are unaffected.
Notes
- Do not stack the
Terse-Coder adapter
on this checkpoint β the preference is already merged in; double
application over-shortens reasoning (63% pass with
no_codefailures). - The fp16 source weights are at Shockem/Qwen3.8-27b-Terse-Coder if you want to quantize differently or merge further.
- Behavioral edit, not a knowledge edit β targeted at coding with thinking enabled. Should work on other backends (SGLang, TabbyAPI/EXL3), but only vLLM has been measured; validate before relying on them.
Attributions & licenses
This checkpoint is a quantized derivative of Shockem/Qwen3.8-27b-Terse-Coder, itself a derivative of Qwen/Qwen3.8-27B, Β© Qwen Team, Alibaba Cloud, licensed Apache 2.0; this checkpoint remains Apache 2.0 and the original license and copyright notices are retained. Credits:
- Qwen Team (Alibaba Cloud) β the Qwen3.8-27B base model (Apache 2.0).
- NVIDIA β TensorRT Model Optimizer 0.45 (Apache 2.0) drove this NVFP4 quantization; NVIDIA's published Qwen3.8-27B-NVFP4 checkpoint informed the Hessian-calibrated recipe.
- agentionai and p-e-w (Heretic) β Signal and a heretic-ara variant were two of the three trace-generation policies in the upstream adapter's preference data.
- OpenAI (HumanEval, MIT) and Google (MBPP, CC-BY 4.0) β prompt sources for training and held-out evaluation.
- Hugging Face TRL (Apache 2.0) β DPO training; Datacurve β DeepSWE, independent evaluation only.
None of these parties endorse this model; all remaining errors are ours.