Swift-Qwen3.8-27B-Uncensored-W4A16-fast

The "fast variant" of Swift-Qwen3.8-27B-Uncensored-W4A16: the same 4-bit body of d0xin/Swift-Qwen3.8-27B-Uncensored-BF16 (UkisAI's reasoning-efficient Swift-Qwen3.8-27B with its refusal direction removed), with the lm_head and the MTP module requantized to int4 GPTQ. It serves on one 24 GB GPU (RTX 3090 class) with the HyperQwen stack. The int4 lm_head makes every decode step read about 0.65 GB less; that is where the gain over the base build comes from.

Lineage

Qwen/Qwen3.8-27B                              Apache 2.0
 └─ ukisai/Swift-Qwen3.8-27b                  Swift Open License v1.0, reasoning-efficiency LoRA merged
     └─ d0xin/Swift-Qwen3.8-27B-Uncensored-BF16   rank-1 directional ablation, layer 38
         └─ ...-W4A16                         W4A16 g128 body, int8 lm_head / embed_tokens / MTP
             └─ this model                    int4 GPTQ lm_head / MTP

What changed against the base build

Part Base build This model
Decoder linear layers int4 g128 symmetric (AutoRound) unchanged (65 of 67 files byte-identical)
lm_head int8 RTN, rel. error 0.0064 int4 GPTQ g128 symmetric, rel. error 0.1407
MTP module (8 linear layers) int8 RTN, rel. error 0.0066–0.0153 int4 GPTQ g128 symmetric, rel. error 0.141–0.170
MTP draft head 40,960 rows of the int8 lm_head 40,960 rows of the int4 lm_head, same id list
embed_tokens, vision tower, norms int8 / BF16 unchanged
Size on disk 16.7 GB 15.8 GB

Relative error is ||dequant - w|| / ||w|| against the BF16 source weights. GPTQ uses 300,000 hidden states captured from this model's own generations, so it minimizes the output error on real activations rather than the weight error that this number shows. On those activations the KL divergence to the BF16 lm_head is 0.00251 for GPTQ int4, against 0.00558 for round-to-nearest int4. All groups are symmetric (no zero points).

The MTP module was evaluated on a held-out split of the captured data (greedy chain simulation, 2 draft positions): 2.269 tokens per step, first-position top-1 agreement 0.777, acceptance 0.845, n = 107,063 positions.

The draft vocabulary is the top 40,960 ids of 6,761 generations by this model (4.59 M output tokens, 3,119 of them with thinking on; prompts: English chat, code, Danish instructions and reasoning, GSM8K train). On the held-out 10 % of those generations it covers 97.68 % of tokens; the generic list that the official Qwen3.8-27B-W4A16-AutoRound checkpoint uses covers 98.18 %. The draft head only matters for SPEC=mtp.

How it was made

  1. drafter/collect_prompts.py, drafter/gen_data.py, drafter/capture.py: generate with this model and capture its final hidden states.
  2. drafter/gptq_lm_head.py --bits 4 --calib-rows 300000: GPTQ int4 lm_head.
  3. drafter/train_mtp.py --eval-only 1 --dump-hessians, then drafter/requant_mtp_gptq.py --bits 4: GPTQ int4 MTP module.
  4. prepare/build_draft_vocab.py: the draft head from the own-output id list.

The tools are in HyperQwen (drafter/, prepare/).

Properties of the BF16 source

These come from the source and were not measured again on this quantization:

  • Refusals (d0xin, fixed 100-prompt set): 88 direct answers, 10 answers with a safety note, 2 other failures, 0 refusals.
  • Capability (d0xin, fixed 298-example set against the original Swift BF16): 39.93 % against 38.26 % combined, +1.68 points, McNemar p = 0.44, no measurable loss.
  • Method: rank-1 directional residual-stream ablation at layer 38; 131 of 1,199 tensors changed, vision tensors unchanged. Details in the source repository.
  • Shorter reasoning (UkisAI): our GSM8K run shows 355 answer tokens on average against 379 for the official AutoRound base, with thinking off.

Measured results (single RTX 3090, 24 GB)

All numbers use the HyperQwen stack on vLLM 0.28.0, WSL2 + Docker, thinking off (protocol v2).

Single-user decode (bench/run_benchmarks.sh single), 2026-09-23, card at 250 W, SPEC=dflash2 (DFlash2 drafter, 7 draft tokens), PREFIX_CACHE=1, KV_MEM=4529848320 (4.22 GiB), MAX_LEN=49152. End-to-end throughput in tok/s, sampling T=default / T=0. One run per checkpoint after a warm-up request.

Concurrency This model Base build Official AutoRound base
C1 123.5 / 121.7 118.0 / 116.7 116.6 / 118.9
C2 175.5 / 190.2 169.5 / 178.5 175.7 / 169.6
C4 242.9 / 248.8 235.8 / 237.1 229.4 / 237.6
C8 214.4 / 254.9 227.5 / 245.7 223.6 / 207.4

C1 accepts 3.67 / 3.75 tokens per verify step, mean time to first token 184 ms. Expect 3–5 % variation between sessions.

Batch serving (bench/run_benchmarks.sh batch --prefill --long), 2026-09-21, card at 350 W, KV=fp8, no speculation, second of two runs.

Row This model Base build
64 concurrent, 128 in / 512 out 1,045.2 tok/s 948.9 tok/s (+10 %)
64 concurrent, 256 in / 256 out 751.7 tok/s 701.4 tok/s (+7 %)
Prefill 1,024 tokens 1,824 tok/s 1,831 tok/s
Prefill 102,400 tokens 1,065 tok/s 1,065 tok/s
1 x 100k prompt: TTFT / TPOT 92.8 s / 24.7 ms 92.8 s / 25.3 ms
4 x 60k prompts, 1,024 out 17.4 tok/s 17.2 tok/s

Quality (bench/quality_battery.py against the served model): perplexity over ~300-token windows of wikitext-2 (en), fineweb-2 Danish (da) and Python source (code); GSM8K exact match, first 200 test questions, greedy, thinking off.

Checkpoint PPL all (en / da / code) GSM8K Mean answer tokens
This model 8.302 (10.85 / 10.98 / 3.35) 97.0 % (same in an earlier run) 355
Base build 8.261 (10.81 / 10.92 / 3.34) 98.0 % (same) 356
Official AutoRound base 8.186 (10.68 / 10.85 / 3.30) 94.5 % 379

The int4 heads cost 0.5 % perplexity. With n=200 the GSM8K standard error is about 1.2 points at this accuracy, so the one-point GSM8K difference is within noise.

How to serve

HyperQwen (Linux or WSL2, one 24 GB GPU), in .env:

MODEL=/app/models/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
SPEC=dflash2          # or SPEC=mtp to draft with the int4 MTP head
PREFIX_CACHE=1

then docker compose --profile single up -d. bash verify.sh --no-server checks the directory before you serve it.

Plain vLLM (0.28 or later) loads the body and heads through compressed-tensors:

vllm serve TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast \
  --max-model-len 32768 --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder

The truncated MTP draft head (mtp.draft_lm_head.* with mtp_draft_vocab_ids.pt) is read by HyperQwen's patches/qwen3_5-mtp-draft-vocab.patch. It is not tested on unpatched vLLM.

Sampling defaults (generation_config.json, from Swift): temperature 1.0, top_p 0.95, top_k 20, min_p 0, repetition_penalty 1.0. For thinking off, Qwen recommends temperature 0.7, top_p 0.8, presence_penalty 1.5. The chat template accepts reasoning_effort low, medium and xhigh (default), and maps minimal to low and high / max to xhigh.

Files

File Content
model-000NN-of-00066.safetensors, model_extra_tensors.safetensors weights (extras: MTP module, draft head; lm_head is in file 66)
model.safetensors.index.json tensor-to-file map
config.json, quantization_config.json architecture and quantization groups (lm_head int4, embed_tokens int8, mtp.* int4)
mtp_draft_vocab_ids.pt, draft_vocab_ids.json the draft head's 40,960 token ids (same list, two formats)
chat_template.jinja Qwen3.8 template with reasoning-effort aliases and string tool arguments
tokenizer.json the Qwen3.8 tokenizer, byte-identical to the source
tokenizer_config.json, generation_config.json, preprocessor_config.json, processor_config.json tokenizer, sampling and vision configuration
LICENSE Swift Open License v1.0 (verbatim from UkisAI)
LICENSE-APACHE-2.0 Apache License 2.0 of the base model (verbatim from Qwen)
NOTICE attribution chain, UkisAI's notices and the list of changes

Safety

This is a refusal-reduced model. It can produce content that the upstream model refuses or handles with more care, and its output can be wrong, offensive or unsafe. Use it for research, development and other lawful purposes. Review its output where an error can cause harm. You are responsible for how you use it and for compliance with the licenses and the law. Fine-tuning should start from the BF16 source, not from these 4-bit weights.

License

The weights are distributed under the Swift Open License v1.0 (LICENSE). It permits use, modification and redistribution. Commercial use is licensed only while you, together with every entity that controls, is controlled by or is under common control with you, had gross revenue below US$1,000,000 in the most recently completed fiscal year (Section 5). Above that threshold, commercial use needs a Swift Enterprise License from UkisAI. The base model's Apache License 2.0 is included as LICENSE-APACHE-2.0 (Section 4(e)); NOTICE lists the attribution chain, UkisAI's notices and every change. This summary is not legal advice; LICENSE is the binding text.

Original Swift model: UkisAI. Abliteration: d0xin. Quantization and serving preparation: TyroneNel.

Downloads last month
95
Safetensors
Model size
4B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast

Base model

Qwen/Qwen3.8-27B
Quantized
(11)
this model