File size: 6,670 Bytes
c3f13a2
 
723d82e
 
e68b1d4
 
 
 
 
 
c3f13a2
e68b1d4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4b53829
 
4742480
 
4b53829
 
 
 
 
e68b1d4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
---
license: apache-2.0
base_model: Shockem/Qwen3.8-27b-Terse-Coder
base_model_relation: quantized
tags:
  - reasoning
  - coding
  - qwen3
  - nvfp4
  - modelopt
---

# Qwen3.8-27B Terse-Coder β€” NVFP4

NVFP4 (modelopt W4A16) quantization of
[Shockem/Qwen3.8-27b-Terse-Coder](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder),
a fine-tune of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)
with **~1/10 the chain-of-thought reasoning tokens on coding tasks and
correctness preserved**. This is the tested deployment artifact β€” every
number below was measured on this checkpoint.

> **Actively researched and improving.** Expect updated quants on this page
> as the study continues.

## Results

Held-out 40 coding problems (20 HumanEval + 20 MBPP-sanitized, disjoint from
training), vLLM 0.28 on 2Γ— RTX 5060 Ti 16 GB, MTP spec decode on, sampling
temp 0.6 / top_k 20 / top_p 0.95 / rep-penalty 1.05, pass@1 by automated
test execution:

| Model (all NVFP4) | pass@1 | Reasoning tokens / problem | Wall tok/s |
|---|---|---|---|
| nvidia/Qwen3.8-27B-NVFP4 (stock) | 72.5% | ~701 | 54.1 |
| **This model** | **67.5%** | **~38 (βˆ’95%)** | **54.5** |

Runs at stock-base wall speed with MTP acceptance 0.412 β€” the reasoning cut
is free end-to-end. **Independent benchmarks** (NVFP4 quant, vLLM 0.28, thinking on, house
sampling; reasoning = `completion_tokens_details.reasoning_tokens`):

| Benchmark | Score | Reasoning tokens (mean / median) |
|---|---|---|
| GSM8K (n=200) | **98.0%** | 84 / 72 |
| GPQA-Diamond (full 198) | **78.3%** | 1,485 / 969 |
| CRUXEval-I (full 800, input prediction) | **92.1%** | 197 / 83 |
| CRUXEval-O (full 800, output prediction) | **92.9%** | 146 / 96 |
| HumanEval+ (164, official EvalPlus, greedy) | **90.2%** (93.9% base) | 43 / 28 |
| MBPP+ (378, official EvalPlus, greedy) | **78.6%** (92.9% base) | 91 / 25 |

CRUXEval was run with the official Meta harness (direct prompts, official
extraction, exec-based scoring, temp 0.2) β€” code *understanding*
(input/output prediction), complementing the generation-side coding table
above.

A note on GPQA-Diamond: this is where a terseness fine-tune is *supposed*
to bleed β€” PhD-level science, far outside the coding training distribution,
where long deliberation is the whole game. Holding **78.3%** at ~1.5k mean
reasoning tokens (thinking models typically burn 10–20k here) means the
training cut the *deliberation budget*, not the *capability* β€” the model
still scales effort up on hard problems (median 969 β†’ max 16k) instead of
answering blindly fast.

**Internal agentic harness** (30 tests across easy/medium/hard β€” instruction
following, coding, reasoning, compaction handoff, tool/JSON contracts β€”
Γ—10 runs each, this checkpoint served by vLLM): **easy 100% (40/40),
medium 100% (90/90), hard 100% (140/140)**, zero truncations, zero
reasoning fallbacks. Prior best on the same harness was 100/100/98.7.

## Quantization recipe

This is a **v3-recipe** house quant, built to preserve the adapter effect
through 4-bit compression:

- modelopt 0.45 **W4A16** NVFP4, per-tensor streaming PTQ (the same
  400-tensor quantize set + ignore list as the published house Signal quants)
- **FP8 attention** (absmax β€” byte-matches NVIDIA's checkpoint at 97–99%)
- **Local-Hessian-weighted calibration on MLP + lm_head** (Hessian captured
  from 2048 house-traffic chunks; Hessian-weighted MSE scale solve with
  per-block e4m3 bracketing). This matters: an absmax-calibrated quant of the
  same weights attenuates the terse-reasoning effect to roughly half
  (βˆ’49.5% vs βˆ’92.4% cut measured). Geomean Hessian-weighted error ratio
  0.805 vs the absmax baseline on the stock base.
- **MTP draft stack included** (1 MTP layer, BF16, vocab-truncated
  40960-id draft head) so speculative decoding works out of the box.

## Serving (vLLM, tested path)

```bash
vllm serve Shockem/Qwen3.8-27b-Terse-Coder-NVFP4 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --kv-cache-dtype fp8
```

**Turn MTP spec decode on** β€” outputs are target-verified (lossless) and
acceptance is 0.41. If you serve with spec decode, make sure the generation
config has **no `min_p`** β€” vLLM 0.28 rejects min_p under spec decode.

Recommended sampling (mirrors testing): temp 0.6, top_k 20, top_p 0.95,
repetition_penalty 1.05.

On 2Γ—16 GB cards cap context at ~200k with a ~3.9 GiB FP8 KV pin;
single-card 24 GB+ rigs are unaffected.

## Notes

- **Do not stack the
  [Terse-Coder adapter](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder-LoRA)
  on this checkpoint** β€” the preference is already merged in; double
  application over-shortens reasoning (63% pass with `no_code` failures).
- The fp16 source weights are at
  [Shockem/Qwen3.8-27b-Terse-Coder](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder)
  if you want to quantize differently or merge further.
- Behavioral edit, not a knowledge edit β€” targeted at coding with thinking
  enabled. Should work on other backends (SGLang, TabbyAPI/EXL3), but only
  vLLM has been measured; validate before relying on them.

## Attributions & licenses

This checkpoint is a quantized derivative of
[Shockem/Qwen3.8-27b-Terse-Coder](https://huggingface.co/Shockem/Qwen3.8-27b-Terse-Coder),
itself a derivative of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B),
Β© Qwen Team, Alibaba Cloud, licensed **Apache 2.0**; this checkpoint remains
Apache 2.0 and the original license and copyright notices are retained.
Credits:

- **Qwen Team (Alibaba Cloud)** β€” the Qwen3.8-27B base model (Apache 2.0).
- **NVIDIA** β€” [TensorRT Model Optimizer](https://github.com/NVIDIA/TensorRT-Model-Optimizer)
  0.45 (Apache 2.0) drove this NVFP4 quantization; NVIDIA's published
  [Qwen3.8-27B-NVFP4](https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4)
  checkpoint informed the Hessian-calibrated recipe.
- **[agentionai](https://huggingface.co/agentionai/Signal-3.8-27B)** and
  **[p-e-w](https://github.com/p-e-w/heretic)** (Heretic) β€” Signal and a
  heretic-ara variant were two of the three trace-generation policies in the
  upstream adapter's preference data.
- **OpenAI** ([HumanEval](https://github.com/openai/human-eval), MIT) and
  **Google** ([MBPP](https://github.com/google-research/google-research/tree/master/mbpp),
  CC-BY 4.0) β€” prompt sources for training and held-out evaluation.
- **Hugging Face [TRL](https://github.com/huggingface/trl)** (Apache 2.0) β€”
  DPO training; **[Datacurve](https://huggingface.co/datasets/datacurve/deep-swe)**
  β€” DeepSWE, independent evaluation only.

None of these parties endorse this model; all remaining errors are ours.