--- base_model: ukisai/Swift-Qwen3.8-27b library_name: transformers license: other license_name: swift-open-license-1.0 pipeline_tag: image-text-to-text tags: - qwen3_8 - efficient-thinking - reasoning - token-efficient - amd - rocm - int4 - awq - quark - w4a16 base_model_relation: quantized ---
# Swift-Qwen3.8-27b-int4-AMD AMD Quark AWQ INT4 (W4A16) edition of Swift. The following introduction describes the base Swift results; release-specific details are below. Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, using **58.3% fewer thinking tokens** while maintaining near-identical performance (**<1% loss**) and as a result getting a **x1.95 speed-up** on several tasks.The prompt is a sample from LiveCodeBench v6
## AMD Quark INT4 release This is the **INT4 W4A16 quantization of Swift for AMD hardware workflows**, produced with [AMD Quark](https://github.com/amd/Quark). It uses Quark's **PyTorch** workflow and native Hugging Face safetensors export: signed symmetric INT4 weights, groups of 128, and BF16 activations. The full-precision companion is [Swift-Qwen3.8-27b-BF16-AMD](https://huggingface.co/ukisai/Swift-Qwen3.8-27b-BF16-AMD). | Property | This checkpoint | | --- | --- | | Source | [Swift-Qwen3.8-27B](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) | | Quantizer | AMD Quark AWQ | | Weight / activation precision | INT4 / BF16 (W4A16) | | Weight grouping | Symmetric, group size 128 | | Format | Native Quark safetensors, `real_quantized`, `reorder` packing | | Weight files | 19.513 GB; BF16 source: 55.563 GB | | Calibration | 128 Pile validation samples, 512 tokens each | | Quantized layers | 496 eligible language-model linear layers | | Preserved components | BF16 vision tower, output head, embeddings, and all 15 MTP tensors | Quark supports preparing models for AMD deployment. **This checkpoint was quantized and validated on an NVIDIA H100; AMD/ROCm serving and throughput have not yet been validated.** Serving needs a runtime that supports this native Quark INT4 format. The [Quark project](https://github.com/amd/Quark) and [installation guide](https://quark.docs.amd.com/latest/install.html) describe its supported CUDA and ROCm environments. ### Checkpoint validation | Sanity check | BF16 | This INT4 export | | --- | ---: | ---: | | Wikitext perplexity | 9.16197 | 9.54254 | | Arithmetic generation | Pass | Pass | | JSON generation | Pass | Pass | Perplexity uses the same eight non-overlapping 512-token Wikitext-2 test windows. The 4.15% perplexity increase is a small sanity result, not a full accuracy benchmark. The packed checkpoint was independently reloaded, including its final configuration and index, and reproduced the evaluation NLLs exactly. All floating tensors are finite; 349 preserved vision/output-head/MTP tensors match the source exactly. Vision inference and MTP decoding were not exercised in this validation. See [quantization_report.json](quantization_report.json). **The Swift benchmarks and speed demonstration below are reproduced from the base Swift model card. They do not measure this Quark export or AMD hardware.** ## Training approach We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in Qwen’s reasoning rollouts. We then fine-tuned Qwen by penalizing usage of those tokens while it reasons. Swift produces shorter reasoning traces. In our testing, we also observe fewer overthinking errors. For maximum gains, Swift also includes a transfer component derived from [BottleCap AI's ThinkingCap-Qwen3.6-27B](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B). ## Evaluation scope > All results below compare the Qwen3.8-27B BF16 base with the same base plus the > Swift adapter. ## Benchmarks| Benchmark | Score | Mean tokens | Median tokens | |||
|---|---|---|---|---|---|---|
| Base | Swift | Base | Swift | Reduction | Reduction | |
| General reasoning | ||||||
| GPQA-Diamond | 88.38% | 88.28% | 15,014 | 8,855 | ↓ 41.0% | ↓ 58.3% |
| MMLU-Pro | 85.47% | 84.95% | 2,980 | 1,603 | ↓ 46.2% | ↓ 28.3% |
| C-Eval | 90.00% | 90.62% | 1,492 | 804 | ↓ 46.1% | ↓ 19.3% |
| IFBench | 73.53% | 71.80% | 8,052 | 4,657 | ↓ 42.2% | ↓ 50.5% |
| Mathematics | ||||||
| AIME 2026 | 98.67% | 94.00% | 22,014 | 16,143 | ↓ 26.7% | ↓ 50.2% |
| HMMT (Nov 2025) | 99.33% | 96.00% | 22,032 | 15,189 | ↓ 31.1% | ↓ 45.9% |
| Multimodal | ||||||
| ERQA | 67.45% | 66.30% | 4,137 | 2,045 | ↓ 50.6% | ↓ 54.6% |
| Agentic coding | ||||||
| Terminal-Bench 2.1 | 66.74% | 65.84% | 37,086 | 27,272 | ↓ 26.5% | ↓ 38.7% |
| LiveCodeBench v6 | 76.76% | 81.55% | 11,374 | 8,615 | ↓ 24.3% | ↓ 45.8% |
Serving: BF16 · vLLM 0.27.1 · Qwen3 parser · context 262,144 · thinking xhigh.
Sampling: temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0 · presence_penalty 0 · repetition_penalty 1.
Benchmarks: averages over five seeds (0–4) per model; five trials per task for Terminal-Bench.
| Benchmark | Output cap |
|---|---|
| GPQA-Diamond | 100,000 |
| MMLU-Pro | 100,000 |
| C-Eval | 16,384 |
| IFBench | 81,920 |
| AIME 2026 | 250,000 |
| HMMT Nov 2025 | 250,000 |
| ERQA | 100,000 |
| Terminal-Bench 2.1 | Agent/task limits |
| LiveCodeBench v6 | 32,768 |
| Reasoning effort | Mean thinking reduction |
|---|---|
| Xhigh | ↓ 41.0% |
| Medium | ↓ 22.7% |
| Low | ↓ 25.8% |
| GPQA-Diamond | Score | Mean tokens | Median tokens |
|---|---|---|---|
| Base · xhigh | 88.38% | 15,014 | 6,642 |
| Swift · xhigh | 88.28% | 8,855 | 2,771 |
| Base · medium | 84.14% | 4,451 | 1,753 |
| Benchmark / quantization | Base accuracy | Swift accuracy | Mean token reduction | Median token reduction |
|---|---|---|---|---|
| GPQA-Diamond Mixed-precision quant W4A16 · thinking tokens | 88.69% | 88.38% | ↓ 32.1% | ↓ 50.2% |
| IFBench Mixed-precision quant W4A16 · completion tokens | 72.58% | 71.25% | ↓ 30.1% | ↓ 38.0% |
| AIME 2026 Mixed-precision quant W4A16 · completion tokens | 84.00% | 84.00% | ↓ 19.0% | ↓ 37.5% |
| AIME 2026 AWQ INT4 · completion tokens | 82.67% | 84.00% | ↓ 22.8% | ↓ 34.8% |