Model card written by GPT-6 Astra, refined with human feedback:

Swift 1.5 HyperQwen — INT4 heads

Swift 1.5's efficient reasoning, combined with HyperQwen's optimized serving, delivered 37% lower average task completion time than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation.

Browse all three Swift HyperQwen variants.

Changes from upstream

Upstream checkpoint: ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ, revision 9dba8a05. The changes below are relative to this AWQ checkpoint.

  • Preserve the upstream AWQ INT4 model body.
  • Convert embeddings to INT8.
  • Quantize the main output head (lm_head) and eight MTP linear matrices to GPTQ INT4, directly from upstream BF16 tensors.
  • Rebuild the MTP draft shortlist using Swift responses: 65,536 tokens, while retaining the target model's full vocabulary.
  • No additional fine-tuning; Swift's reasoning-efficiency training is retained.

Performance

Measurement Qwen — HyperQwen fast quant (W4A16 AutoRound) Swift 1.0 Swift 1.5 INT8 heads Swift 1.5 INT4 heads
Average request time ↓ 108.1 s 66.2 s 72.2 s 68.2 s
Average output tokens/task ↓ 8,985 5,245 5,751 5,669
Median decode tokens/s ↑ 112.1 105.9 104.0 107.2
Total output tokens, 630 tasks 5.66M 3.30M 3.62M 3.57M

Quality

Test Qwen — HyperQwen fast quant (W4A16 AutoRound) Swift 1.0 Swift 1.5 INT8 heads Swift 1.5 INT4 heads
GSM8K — 200-question subset 97.5% 98.0% 98.0% 97.5%
IFBench — 300 prompts, strict 74.0% 73.3% 73.7% 72.3%
LiveCodeBench — 100-problem subset 90% 89% 89% 91%
Custom tool-call/JSON checks — 30 tasks 29/30 28/30 30/30 30/30
Perplexity, English/Python ↓ 6.551 6.605 6.643 6.679
Truncated answers, counted wrong 2 1 1 0
  • GSM8K: the first 200 questions from the test split, with thinking disabled.
  • IFBench: all 300 prompts in the pinned IFBench test dataset, scored with the official strict prompt-level verifier.
  • LiveCodeBench: a frozen v6-era dataset subset of Python stdin/stdout problems: 34 easy, 33 medium, 33 hard. Scored against supplied public/private tests with a custom judge; not a full official LiveCodeBench result.
  • Tool-call/JSON checks: 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark.
  • Perplexity: 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see subset results.

Compared with the Swift 1.5 INT8-head variant, this INT4-head conversion measured 3.2% higher decode TPS and 5.5% lower average request time; benchmark score changes were mixed.

Evaluation setup: RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy.

Setup

Requires the patched HyperQwen runtime, not stock vLLM or GGUF tools. From an installed HyperQwen checkout:

hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" CTX=long MAX_LEN=150000 SPEC=mtp \
  bash single-user/start_qwen.sh

See RUNTIME.md for the pinned runtime, launcher snapshot, complete evaluated settings and installation notes. The checkpoint is already converted: do not requantize its heads. Vision weights are retained, but the evaluation is text-only. Multi-user batch settings were not benchmarked in this campaign.

Detailed results and evaluation code are in evaluation/. Quantization/source provenance is included with the model. The upstream Swift Open License v1.0, Apache 2.0 base-model license, and NOTICE are retained.

The upstream AWQ quantization is UkisAI's Swift 1.5 W4A16-AWQ. Swift's training is by UkisAI; the serving runtime is HyperQwen. This is an independent conversion with local evaluation by daavidhauser.

Downloads last month
-
Safetensors
Model size
28B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4

Base model

Qwen/Qwen3.8-27B
Quantized
(34)
this model

Collection including daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4