Model card written by GPT-6 Astra, refined with human feedback:
Swift 1.5 HyperQwen — INT4 heads
Swift 1.5's efficient reasoning, combined with HyperQwen's optimized serving, delivered 37% lower average task completion time than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation.
Browse all three Swift HyperQwen variants.
Changes from upstream
Upstream checkpoint: ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ, revision 9dba8a05. The changes below are relative to this AWQ checkpoint.
- Preserve the upstream AWQ INT4 model body.
- Convert embeddings to INT8.
- Quantize the main output head (
lm_head) and eight MTP linear matrices to GPTQ INT4, directly from upstream BF16 tensors. - Rebuild the MTP draft shortlist using Swift responses: 65,536 tokens, while retaining the target model's full vocabulary.
- No additional fine-tuning; Swift's reasoning-efficiency training is retained.
Performance
| Measurement | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads |
|---|---|---|---|---|
| Average request time ↓ | 108.1 s | 66.2 s | 72.2 s | 68.2 s |
| Average output tokens/task ↓ | 8,985 | 5,245 | 5,751 | 5,669 |
| Median decode tokens/s ↑ | 112.1 | 105.9 | 104.0 | 107.2 |
| Total output tokens, 630 tasks | 5.66M | 3.30M | 3.62M | 3.57M |
Quality
| Test | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads |
|---|---|---|---|---|
| GSM8K — 200-question subset | 97.5% | 98.0% | 98.0% | 97.5% |
| IFBench — 300 prompts, strict | 74.0% | 73.3% | 73.7% | 72.3% |
| LiveCodeBench — 100-problem subset | 90% | 89% | 89% | 91% |
| Custom tool-call/JSON checks — 30 tasks | 29/30 | 28/30 | 30/30 | 30/30 |
| Perplexity, English/Python ↓ | 6.551 | 6.605 | 6.643 | 6.679 |
| Truncated answers, counted wrong | 2 | 1 | 1 | 0 |
- GSM8K: the first 200 questions from the test split, with thinking disabled.
- IFBench: all 300 prompts in the pinned IFBench test dataset, scored with the official strict prompt-level verifier.
- LiveCodeBench: a frozen v6-era dataset subset of Python stdin/stdout problems: 34 easy, 33 medium, 33 hard. Scored against supplied public/private tests with a custom judge; not a full official LiveCodeBench result.
- Tool-call/JSON checks: 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark.
- Perplexity: 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see subset results.
Compared with the Swift 1.5 INT8-head variant, this INT4-head conversion measured 3.2% higher decode TPS and 5.5% lower average request time; benchmark score changes were mixed.
Evaluation setup: RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy.
Setup
Requires the patched HyperQwen runtime, not stock vLLM or GGUF tools. From an installed HyperQwen checkout:
hf download daavidhauser/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4 --local-dir models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4
MODEL="$PWD/models/Swift-1.5-Qwen3.8-27B-W4A16-HyperQwen-INT4" CTX=long MAX_LEN=150000 SPEC=mtp \
bash single-user/start_qwen.sh
See RUNTIME.md for the pinned runtime, launcher snapshot, complete evaluated settings and installation notes. The checkpoint is already converted: do not requantize its heads. Vision weights are retained, but the evaluation is text-only. Multi-user batch settings were not benchmarked in this campaign.
Detailed results and evaluation code are in evaluation/. Quantization/source provenance is included with the model. The upstream Swift Open License v1.0, Apache 2.0 base-model license, and NOTICE are retained.
The upstream AWQ quantization is UkisAI's Swift 1.5 W4A16-AWQ. Swift's training is by UkisAI; the serving runtime is HyperQwen. This is an independent conversion with local evaluation by daavidhauser.
- Downloads last month
- -