--- license: other license_name: swift-open-license-1.0 license_link: LICENSE base_model: ukisai/Swift-Qwen3.8-27b base_model_relation: quantized library_name: vllm pipeline_tag: text-generation tags: - compressed-tensors - awq - hyperqwen - efficient-thinking - int8-heads --- Model card written by GPT-6 Astra, refined with human feedback: # Swift 1.0 HyperQwen Swift 1.0's efficient reasoning, combined with HyperQwen's optimized serving, **delivered 39% lower average task completion time** than HyperQwen serving the Qwen fast checkpoint in our RTX 3090 evaluation. This checkpoint brings Swift 1.0 to HyperQwen's quantized serving and MTP speculative-decoding path. Across 630 tasks, it generated **42% fewer output tokens** than the Qwen fast reference. Swift reduces the amount of generation needed to finish a task; HyperQwen accelerates generation. The headline result is shorter task time, even where raw tokens per second is lower. Average task time here is per-request elapsed time at **two concurrent requests**. [Browse all three Swift HyperQwen variants](https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks-6abac232566f58d5e7c5a046). ## Changes from upstream Upstream checkpoint: [TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ](https://huggingface.co/TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ), revision [`6ced337b`](https://huggingface.co/TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ/tree/6ced337b9c99adddf9871ff84abc78a7f1acafad). The changes below are relative to this AWQ checkpoint. - Preserve the upstream **AWQ INT4 model body**. - Convert embeddings to **INT8**. - Convert the main output head (`lm_head`) and MTP linear weights to **INT8**. - Add HyperQwen's reference MTP draft shortlist, while retaining the target model's full vocabulary. - No additional fine-tuning; Swift's reasoning-efficiency training is retained. ## Performance | Measurement | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads | |---|---:|---:|---:|---:| | **Average request time ↓** | **108.1 s** | **66.2 s** | **72.2 s** | **68.2 s** | | **Average output tokens/task ↓** | **8,985** | **5,245** | **5,751** | **5,669** | | Median decode tokens/s ↑ | 112.1 | 105.9 | 104.0 | 107.2 | | Total output tokens, 630 tasks | 5.66M | 3.30M | 3.62M | 3.57M | ## Quality | Test | Qwen — HyperQwen fast quant (W4A16 AutoRound) | Swift 1.0 | Swift 1.5 INT8 heads | Swift 1.5 INT4 heads | |---|---:|---:|---:|---:| | **GSM8K — 200-question subset** | 97.5% | 98.0% | 98.0% | 97.5% | | **IFBench — 300 prompts, strict** | 74.0% | 73.3% | 73.7% | 72.3% | | **LiveCodeBench — 100-problem subset** | 90% | 89% | 89% | 91% | | **Custom tool-call/JSON checks — 30 tasks** | 29/30 | 28/30 | 30/30 | 30/30 | | Perplexity, English/Python ↓ | 6.551 | 6.605 | 6.643 | 6.679 | | Truncated answers, counted wrong | 2 | 1 | 1 | 0 | - **GSM8K:** the first 200 questions from the [test split](https://huggingface.co/datasets/openai/gsm8k), with thinking disabled. - **IFBench:** all 300 prompts in the pinned [IFBench test dataset](https://huggingface.co/datasets/allenai/IFBench_test), scored with the official strict prompt-level verifier. - **LiveCodeBench:** a frozen [v6-era dataset](https://huggingface.co/datasets/livecodebench/code_generation_lite) subset of Python stdin/stdout problems: 34 easy, 33 medium, 33 hard. Scored against supplied public/private tests with a custom judge; not a full official LiveCodeBench result. - **Tool-call/JSON checks:** 20 custom weather-tool tasks checking the function name, arguments, Celsius-to-Fahrenheit conversion and final JSON; plus 10 JSON inventory-filtering tasks. These are integration checks, not an external agent benchmark. - **Perplexity:** 18,729 scored tokens from English Wikipedia and Python source; lower is better. Recomputed from saved token log-probabilities, excluding Danish; see [subset results](evaluation/perplexity-english-python.json). **Evaluation setup:** RTX 3090 24 GB; FP8 KV cache; 150,000-token configured context; 128,000 output tokens per call. These are runtime settings, not fixed model properties, and the TPS test uses short prompts. Task evaluation used two concurrent requests; TPS was measured with one request at a time. All models used the same serving settings and task budgets. Thinking tests used xhigh effort, temperature 1.0, top_p 0.95, top_k 20 and seed 15027; GSM8K/tool checks were greedy. ## Setup Requires the **patched HyperQwen runtime**, not stock vLLM or GGUF tools. From an installed HyperQwen checkout: ```bash hf download daavidhauser/Swift-1.0-Qwen3.8-27B-W4A16-HyperQwen --local-dir models/Swift-1.0-Qwen3.8-27B-W4A16-HyperQwen MODEL="$PWD/models/Swift-1.0-Qwen3.8-27B-W4A16-HyperQwen" CTX=long MAX_LEN=150000 SPEC=mtp \ bash single-user/start_qwen.sh ``` See [RUNTIME.md](RUNTIME.md) for the pinned runtime, launcher snapshot, complete evaluated settings and installation notes. The checkpoint is already converted: do not requantize its heads. Vision weights are retained, but the evaluation is text-only. Multi-user batch settings were not benchmarked in this campaign. Detailed results and evaluation code are in [evaluation/](evaluation/). Quantization/source provenance is included with the model. The upstream [Swift Open License v1.0](LICENSE), [Apache 2.0 base-model license](LICENSE-APACHE-2.0), and [NOTICE](NOTICE) are retained. The upstream AWQ quantization is [TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ](https://huggingface.co/TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ). Swift's training is by [UkisAI](https://huggingface.co/ukisai); the serving runtime is [HyperQwen](https://github.com/syv-ai/HyperQwen). This is an independent conversion with local evaluation by daavidhauser.