File size: 6,670 Bytes
c3f13a2 723d82e e68b1d4 c3f13a2 e68b1d4 4b53829 4742480 4b53829 e68b1d4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | ---
license: apache-2.0
base_model: Shockem/Qwen3.8-27b-Terse-Coder
base_model_relation: quantized
tags:
- reasoning
- coding
- qwen3
- nvfp4
- modelopt
---
# Qwen3.8-27B Terse-Coder β NVFP4
NVFP4 (modelopt W4A16) quantization of
[Shockem/Qwen3.8-27b-Terse-Coder](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder),
a fine-tune of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)
with **~1/10 the chain-of-thought reasoning tokens on coding tasks and
correctness preserved**. This is the tested deployment artifact β every
number below was measured on this checkpoint.
> **Actively researched and improving.** Expect updated quants on this page
> as the study continues.
## Results
Held-out 40 coding problems (20 HumanEval + 20 MBPP-sanitized, disjoint from
training), vLLM 0.28 on 2Γ RTX 5060 Ti 16 GB, MTP spec decode on, sampling
temp 0.6 / top_k 20 / top_p 0.95 / rep-penalty 1.05, pass@1 by automated
test execution:
| Model (all NVFP4) | pass@1 | Reasoning tokens / problem | Wall tok/s |
|---|---|---|---|
| nvidia/Qwen3.8-27B-NVFP4 (stock) | 72.5% | ~701 | 54.1 |
| **This model** | **67.5%** | **~38 (β95%)** | **54.5** |
Runs at stock-base wall speed with MTP acceptance 0.412 β the reasoning cut
is free end-to-end. **Independent benchmarks** (NVFP4 quant, vLLM 0.28, thinking on, house
sampling; reasoning = `completion_tokens_details.reasoning_tokens`):
| Benchmark | Score | Reasoning tokens (mean / median) |
|---|---|---|
| GSM8K (n=200) | **98.0%** | 84 / 72 |
| GPQA-Diamond (full 198) | **78.3%** | 1,485 / 969 |
| CRUXEval-I (full 800, input prediction) | **92.1%** | 197 / 83 |
| CRUXEval-O (full 800, output prediction) | **92.9%** | 146 / 96 |
| HumanEval+ (164, official EvalPlus, greedy) | **90.2%** (93.9% base) | 43 / 28 |
| MBPP+ (378, official EvalPlus, greedy) | **78.6%** (92.9% base) | 91 / 25 |
CRUXEval was run with the official Meta harness (direct prompts, official
extraction, exec-based scoring, temp 0.2) β code *understanding*
(input/output prediction), complementing the generation-side coding table
above.
A note on GPQA-Diamond: this is where a terseness fine-tune is *supposed*
to bleed β PhD-level science, far outside the coding training distribution,
where long deliberation is the whole game. Holding **78.3%** at ~1.5k mean
reasoning tokens (thinking models typically burn 10β20k here) means the
training cut the *deliberation budget*, not the *capability* β the model
still scales effort up on hard problems (median 969 β max 16k) instead of
answering blindly fast.
**Internal agentic harness** (30 tests across easy/medium/hard β instruction
following, coding, reasoning, compaction handoff, tool/JSON contracts β
Γ10 runs each, this checkpoint served by vLLM): **easy 100% (40/40),
medium 100% (90/90), hard 100% (140/140)**, zero truncations, zero
reasoning fallbacks. Prior best on the same harness was 100/100/98.7.
## Quantization recipe
This is a **v3-recipe** house quant, built to preserve the adapter effect
through 4-bit compression:
- modelopt 0.45 **W4A16** NVFP4, per-tensor streaming PTQ (the same
400-tensor quantize set + ignore list as the published house Signal quants)
- **FP8 attention** (absmax β byte-matches NVIDIA's checkpoint at 97β99%)
- **Local-Hessian-weighted calibration on MLP + lm_head** (Hessian captured
from 2048 house-traffic chunks; Hessian-weighted MSE scale solve with
per-block e4m3 bracketing). This matters: an absmax-calibrated quant of the
same weights attenuates the terse-reasoning effect to roughly half
(β49.5% vs β92.4% cut measured). Geomean Hessian-weighted error ratio
0.805 vs the absmax baseline on the stock base.
- **MTP draft stack included** (1 MTP layer, BF16, vocab-truncated
40960-id draft head) so speculative decoding works out of the box.
## Serving (vLLM, tested path)
```bash
vllm serve Shockem/Qwen3.8-27b-Terse-Coder-NVFP4 \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--kv-cache-dtype fp8
```
**Turn MTP spec decode on** β outputs are target-verified (lossless) and
acceptance is 0.41. If you serve with spec decode, make sure the generation
config has **no `min_p`** β vLLM 0.28 rejects min_p under spec decode.
Recommended sampling (mirrors testing): temp 0.6, top_k 20, top_p 0.95,
repetition_penalty 1.05.
On 2Γ16 GB cards cap context at ~200k with a ~3.9 GiB FP8 KV pin;
single-card 24 GB+ rigs are unaffected.
## Notes
- **Do not stack the
[Terse-Coder adapter](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder-LoRA)
on this checkpoint** β the preference is already merged in; double
application over-shortens reasoning (63% pass with `no_code` failures).
- The fp16 source weights are at
[Shockem/Qwen3.8-27b-Terse-Coder](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder)
if you want to quantize differently or merge further.
- Behavioral edit, not a knowledge edit β targeted at coding with thinking
enabled. Should work on other backends (SGLang, TabbyAPI/EXL3), but only
vLLM has been measured; validate before relying on them.
## Attributions & licenses
This checkpoint is a quantized derivative of
[Shockem/Qwen3.8-27b-Terse-Coder](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder),
itself a derivative of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B),
Β© Qwen Team, Alibaba Cloud, licensed **Apache 2.0**; this checkpoint remains
Apache 2.0 and the original license and copyright notices are retained.
Credits:
- **Qwen Team (Alibaba Cloud)** β the Qwen3.8-27B base model (Apache 2.0).
- **NVIDIA** β [TensorRT Model Optimizer](https://github.com/NVIDIA/TensorRT-Model-Optimizer)
0.45 (Apache 2.0) drove this NVFP4 quantization; NVIDIA's published
[Qwen3.8-27B-NVFP4](https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4)
checkpoint informed the Hessian-calibrated recipe.
- **[agentionai](https://huggingface.co/agentionai/Signal-3.8-27B)** and
**[p-e-w](https://github.com/p-e-w/heretic)** (Heretic) β Signal and a
heretic-ara variant were two of the three trace-generation policies in the
upstream adapter's preference data.
- **OpenAI** ([HumanEval](https://github.com/openai/human-eval), MIT) and
**Google** ([MBPP](https://github.com/google-research/google-research/tree/master/mbpp),
CC-BY 4.0) β prompt sources for training and held-out evaluation.
- **Hugging Face [TRL](https://github.com/huggingface/trl)** (Apache 2.0) β
DPO training; **[Datacurve](https://huggingface.co/datasets/datacurve/deep-swe)**
β DeepSWE, independent evaluation only.
None of these parties endorse this model; all remaining errors are ours.
|