Image-Text-to-Text
Transformers
Safetensors
qwen3_5
paroquant
mxfp6
w6a8
rocm
rdna4
conversational
6-bit
paroquant_mxfp6
Instructions to use realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6") model = AutoModelForMultimodalLM.from_pretrained("realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6
- SGLang
How to use realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6 with Docker Model Runner:
docker model run hf.co/realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6
Model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,398 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: swift-open-license-1.0
|
| 4 |
+
license_link: LICENSE
|
| 5 |
+
library_name: transformers
|
| 6 |
+
pipeline_tag: image-text-to-text
|
| 7 |
+
base_model: ukisai/Swift-1.5-Qwen3.8-27b
|
| 8 |
+
base_model_relation: quantized
|
| 9 |
+
tags:
|
| 10 |
+
- paroquant
|
| 11 |
+
- mxfp6
|
| 12 |
+
- w6a8
|
| 13 |
+
- rocm
|
| 14 |
+
- rdna4
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# Swift-1.5-Qwen3.8-27B-PARO-MXFP6
|
| 18 |
+
|
| 19 |
+
ParoQuant MXFP6 quant of [`ukisai/Swift-1.5-Qwen3.8-27b`](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b). Unofficial.
|
| 20 |
+
|
| 21 |
+
- Scheme: MXFP6 E2M3 weights (W6A8), ParoQuant rotations (krot 8, group 128), `quant_method: paroquant_mxfp6`
|
| 22 |
+
- Method: rotations from [`z-lab/Qwen3.8-27B-PARO`](https://huggingface.co/z-lab/Qwen3.8-27B-PARO) (trained on base Qwen3.8-27B, frozen), then round-to-nearest to MXFP6. Stage-2 fine-tune skipped for time.
|
| 23 |
+
- Same recipe as [`hugypufy/Swift-Qwen3.8-27B-PARO-MXFP6`](https://huggingface.co/hugypufy/Swift-Qwen3.8-27B-PARO-MXFP6), minus its fine-tune
|
| 24 |
+
- Needs a vLLM build with the `paroquant_mxfp6` method (ROCm / RDNA4)
|
| 25 |
+
- ~24 GB
|
| 26 |
+
|
| 27 |
+
---
|
| 28 |
+
|
| 29 |
+
# Original model card
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
<div align="center">
|
| 33 |
+
<a href="https://ukisai.com"><img src="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/resolve/main/ukisai-banner.png" alt="UkisAI" style="width:100%;max-width:100%;height:auto;display:block;margin-bottom:0.6em;" /></a>
|
| 34 |
+
<div style="display:flex;justify-content:center;gap:0.6em;margin-bottom:1em;">
|
| 35 |
+
<a href="https://ukisai.com"><strong>Website</strong></a> •
|
| 36 |
+
<a href="https://ukisai.com/products/swift"><strong>Learn more</strong></a> •
|
| 37 |
+
<a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GGUF"><strong>GGUF</strong></a> •
|
| 38 |
+
<a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF"><strong>GSQ-RCO GGUF</strong></a> •
|
| 39 |
+
<a href="#evaluation"><strong>Evaluation</strong></a> •
|
| 40 |
+
<a href="#license-and-access"><strong>Enterprise licensing</strong></a>
|
| 41 |
+
</div>
|
| 42 |
+
</div>
|
| 43 |
+
|
| 44 |
+
# Swift 1.5 Qwen3.8-27B
|
| 45 |
+
|
| 46 |
+
Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient derivative of [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B).
|
| 47 |
+
It uses **58.5% fewer thinking tokens** while scoring **0.35% higher** than the base, for a **1.95× speed-up** on several tasks.
|
| 48 |
+
|
| 49 |
+
Swift 1.5 is a direct upgrade from [Swift 1.0](https://huggingface.co/ukisai/Swift-Qwen3.8-27b), our model with 350k+ downloads, delivering stronger overall performance than both base and Swift 1.0 in various tasks, especially coding and agentic, while using fewer thinking tokens. We accomplished that by scaling up the post-training (RL and OPD) from the previous version.
|
| 50 |
+
|
| 51 |
+
## Demo
|
| 52 |
+
|
| 53 |
+
We gave base Qwen3.8-27B and Swift 1.5 27B the same prompt:
|
| 54 |
+
|
| 55 |
+
> create a 3d little planet globe where I (player can walk around) and it has all these biomes to explore, the globe doesn't have to be too big, but still fun to go around. It's about a boy scout who is camping and goes around exploring.
|
| 56 |
+
|
| 57 |
+
<video src="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/resolve/main/swift-1.5-planet-demo.mp4" controls autoplay muted loop playsinline style="width:100%;height:auto;border-radius:12px;"></video>
|
| 58 |
+
|
| 59 |
+
Try the game yourself here: [https://ukisai.com/swift-games/27b](https://ukisai.com/swift-games/27b)
|
| 60 |
+
|
| 61 |
+
Base Qwen3.8-27B took 104.6 minutes to build its game. Swift 1.5 took 11.39 minutes.
|
| 62 |
+
|
| 63 |
+
## Training approach
|
| 64 |
+
|
| 65 |
+
We made Swift efficient by figuring out which tokens were linked to pathological overthinking and penalizing them without "attacking" the reasoning length directly then regained the accuracy with RL and OPD, leading to "compressed" token usage while maintaining accuracy.
|
| 66 |
+
Swift 1.5 was made from [Swift 1.0](https://huggingface.co/ukisai/Swift-Qwen3.8-27b), on whom we scaled up the post-training methods that previously improved Swift1.0 model performance, this time with the main
|
| 67 |
+
focus on long-horizon, agentic, and coding tasks, as seen in the LiveCodeBench and Terminal Bench 2.1 improvements. Our training data is viewable here: https://huggingface.co/datasets/ukisai/Qwen3.8-27B-multi-turn-agent-sft albeit it is not used out of the box, but rather re-sampled, turned into proper RL environments etc.
|
| 68 |
+
|
| 69 |
+
## Evaluation
|
| 70 |
+
|
| 71 |
+
The external results below compare **Qwen3.8-27B**, the foundation base model,
|
| 72 |
+
and **Swift 1.5**. Both models use
|
| 73 |
+
the same saved evaluation protocols, and all scores are reported as final aggregate
|
| 74 |
+
percentages.
|
| 75 |
+
|
| 76 |
+
<style>
|
| 77 |
+
.swift15-table { width:100%; table-layout:fixed; border-collapse:separate; border-spacing:0; overflow:hidden; border:1px solid #27344A; border-radius:20px; background:#0D111B; font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,sans-serif; font-size:14px; color:#BFBDBD; }
|
| 78 |
+
.swift15-table th { padding:13px 8px; text-align:center; font-weight:700; color:#AEB5C7; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; }
|
| 79 |
+
.swift15-table td { padding:14px 8px; text-align:center; color:#BFBDBD; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; vertical-align:middle; overflow-wrap:break-word; }
|
| 80 |
+
.swift15-table tr > :last-child { border-right:0; }
|
| 81 |
+
.swift15-table tbody tr:last-child td { border-bottom:0; }
|
| 82 |
+
.swift15-table .benchmark-heading { color:#B7BDCD; background:#0D111B; border-bottom:3px solid #7D45B5; }
|
| 83 |
+
.swift15-table .score-heading { color:#F0C5FF; background:#52239E; border-bottom:3px solid #7D45B5; }
|
| 84 |
+
.swift15-table .tokens-heading, .swift15-table .median-heading { color:#D4E8FF; background:#304FC2; border-bottom:3px solid #5687E6; }
|
| 85 |
+
.swift15-table .change-heading { color:#61B9FF; background:#101D2D; border-bottom:3px solid #5687E6; }
|
| 86 |
+
.swift15-table .change { color:#61B9FF; background:#101D2D; font-weight:700; }
|
| 87 |
+
.swift15-table .gain { color:#DDA8FF; background:#171127; font-weight:700; }
|
| 88 |
+
.swift15-table .model { padding-left:18px; text-align:left; color:#FFFFFF; font-weight:600; }
|
| 89 |
+
.swift15-table strong { color:#FFFFFF; }
|
| 90 |
+
.swift15-table .section { padding:12px 18px; text-align:left; color:#B489FF; background:#2A2541; font-weight:700; letter-spacing:.08em; text-transform:uppercase; border-top:1px solid #3A3159; border-bottom:1px solid #3A3159; }
|
| 91 |
+
.swift15-table .swift15 { background:#171127; }
|
| 92 |
+
.swift15-table thead tr:nth-child(2) .swift15 { color:#D3A0FF; }
|
| 93 |
+
.swift15-table .detail { color:#8C94A8; font-size:12px; font-weight:500; }
|
| 94 |
+
.swift15-table.swift15-compact { display:table !important; width:100% !important; table-layout:fixed !important; }
|
| 95 |
+
|
| 96 |
+
@media (max-width: 640px) {
|
| 97 |
+
.swift15-table { display:block !important; width:100% !important; max-width:100%; overflow-x:auto !important; -webkit-overflow-scrolling:touch; table-layout:auto !important; }
|
| 98 |
+
.swift15-table th, .swift15-table td { min-width:104px; }
|
| 99 |
+
.swift15-table th:first-child, .swift15-table td:first-child { min-width:170px; }
|
| 100 |
+
}
|
| 101 |
+
</style>
|
| 102 |
+
|
| 103 |
+
<table class="swift15-table">
|
| 104 |
+
<thead>
|
| 105 |
+
<tr>
|
| 106 |
+
<th class="benchmark-heading" rowspan="2" style="width:32%;text-align:left;padding-left:18px;vertical-align:bottom;">Benchmark</th>
|
| 107 |
+
<th class="score-heading" colspan="2">Final score</th>
|
| 108 |
+
<th class="tokens-heading" colspan="3">Mean tokens</th>
|
| 109 |
+
<th class="median-heading">Median tokens</th>
|
| 110 |
+
</tr>
|
| 111 |
+
<tr>
|
| 112 |
+
<th>Qwen3.8</th>
|
| 113 |
+
<th class="swift15">Swift 1.5</th>
|
| 114 |
+
<th>Qwen3.8</th>
|
| 115 |
+
<th class="swift15">Swift 1.5</th>
|
| 116 |
+
<th class="change-heading">Reduction</th>
|
| 117 |
+
<th class="change-heading">Reduction</th>
|
| 118 |
+
</tr>
|
| 119 |
+
</thead>
|
| 120 |
+
<tbody>
|
| 121 |
+
<tr><td class="section" colspan="7">General reasoning</td></tr>
|
| 122 |
+
<tr><td class="model">GPQA-Diamond</td><td>88.28%</td><td class="swift15"><strong>88.59%</strong></td><td>15,014</td><td class="swift15">8,717</td><td class="change">↓ 41.9%</td><td class="change">↓ 58.5%</td></tr>
|
| 123 |
+
<tr><td class="model">C-Eval</td><td>90.00%</td><td class="swift15"><strong>90.92%</strong></td><td>1,492</td><td class="swift15">819</td><td class="change">↓ 45.1%</td><td class="change">↓ 16.9%</td></tr>
|
| 124 |
+
<tr><td class="model">IFBench</td><td><strong>73.53%</strong></td><td class="swift15">72.07%</td><td>8,052</td><td class="swift15">4,955</td><td class="change">↓ 38.5%</td><td class="change">↓ 47.3%</td></tr>
|
| 125 |
+
<tr><td class="model">ERQA</td><td><strong>67.45%</strong></td><td class="swift15">65.40%</td><td>4,137</td><td class="swift15">1,906</td><td class="change">↓ 53.9%</td><td class="change">↓ 56.2%</td></tr>
|
| 126 |
+
<tr><td class="section" colspan="7">Mathematics</td></tr>
|
| 127 |
+
<tr><td class="model">AIME 2026</td><td><strong>98.67%</strong></td><td class="swift15">96.00%</td><td>22,014</td><td class="swift15">13,203</td><td class="change">↓ 40.0%</td><td class="change">↓ 48.5%</td></tr>
|
| 128 |
+
<tr><td class="model">HMMT November 2025</td><td><strong>99.33%</strong></td><td class="swift15">97.33%</td><td>22,032</td><td class="swift15">14,957</td><td class="change">↓ 32.1%</td><td class="change">↓ 47.8%</td></tr>
|
| 129 |
+
<tr><td class="section" colspan="7">Coding</td></tr>
|
| 130 |
+
<tr><td class="model">LiveCodeBench v6</td><td>76.76%</td><td class="swift15"><strong>81.71%</strong></td><td>11,184</td><td class="swift15">8,448</td><td class="change">↓ 24.5%</td><td class="change">↓ 46.3%</td></tr>
|
| 131 |
+
<tr><td class="section" colspan="7">Agent tasks</td></tr>
|
| 132 |
+
<tr><td class="model">Terminal-Bench 2.1*</td><td>69.21%</td><td class="swift15"><strong>72.13%</strong></td><td>52,265</td><td class="swift15">43,733</td><td class="change">↓ 16.3%</td><td class="change">↓ 0.1%</td></tr>
|
| 133 |
+
</tbody>
|
| 134 |
+
</table>
|
| 135 |
+
|
| 136 |
+
<p><strong>* Note:</strong> Terminal Bench 2.1 score of Swift1.5 27B is misleadingly low at first glance. It is not a bug, but a simple matter of the Swift models not falling into overthinking loops and failing the task, rather pursuing it until the end, leading to higher average token usage. The token reduction still falls in the -38.7% range when compared apples-to-apples.</p>
|
| 137 |
+
|
| 138 |
+
<details>
|
| 139 |
+
<summary><strong>Benchmark methodology and reproduction settings</strong></summary>
|
| 140 |
+
|
| 141 |
+
<p style="font-size:13px;line-height:1.5;margin:8px 0;"><strong>Serving:</strong> BF16 · vLLM 0.27.1 · Qwen3 parser · context 262,144 · thinking xhigh.<br>
|
| 142 |
+
<strong>Sampling:</strong> temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0 · presence_penalty 0 · repetition_penalty 1.<br>
|
| 143 |
+
<strong>Benchmarks:</strong> averages over five seeds (0–4) per model; five trials per task for Terminal-Bench, base and Swift 1.5 served at context 131,072 on the same Harbor build.</p>
|
| 144 |
+
|
| 145 |
+
<table style="display:table;width:100%;border-collapse:collapse;font-size:13px;line-height:1.3;margin:8px 0;">
|
| 146 |
+
<thead><tr><th style="padding:4px 8px;text-align:left;">Benchmark</th><th style="padding:4px 8px;text-align:right;">Output cap</th></tr></thead>
|
| 147 |
+
<tbody>
|
| 148 |
+
<tr><td style="padding:3px 8px;">GPQA-Diamond</td><td style="padding:3px 8px;text-align:right;">100,000</td></tr>
|
| 149 |
+
<tr><td style="padding:3px 8px;">C-Eval</td><td style="padding:3px 8px;text-align:right;">16,384</td></tr>
|
| 150 |
+
<tr><td style="padding:3px 8px;">IFBench</td><td style="padding:3px 8px;text-align:right;">81,920</td></tr>
|
| 151 |
+
<tr><td style="padding:3px 8px;">ERQA</td><td style="padding:3px 8px;text-align:right;">100,000</td></tr>
|
| 152 |
+
<tr><td style="padding:3px 8px;">AIME 2026</td><td style="padding:3px 8px;text-align:right;">250,000</td></tr>
|
| 153 |
+
<tr><td style="padding:3px 8px;">HMMT November 2025</td><td style="padding:3px 8px;text-align:right;">250,000</td></tr>
|
| 154 |
+
<tr><td style="padding:3px 8px;">LiveCodeBench v6</td><td style="padding:3px 8px;text-align:right;">32,768</td></tr>
|
| 155 |
+
<tr><td style="padding:3px 8px;">Terminal-Bench 2.1</td><td style="padding:3px 8px;text-align:right;">Agent/task limits</td></tr>
|
| 156 |
+
</tbody>
|
| 157 |
+
</table>
|
| 158 |
+
|
| 159 |
+
</details>
|
| 160 |
+
|
| 161 |
+
## Efficiency across reasoning efforts
|
| 162 |
+
|
| 163 |
+
Qwen3.8's `reasoning_effort` setting lets users choose how much the model thinks.
|
| 164 |
+
For Swift 1.5 to be useful across these settings, it needs to reduce thinking while
|
| 165 |
+
keeping accuracy close to the base. We therefore tested `xhigh`, `medium`, and `low`:
|
| 166 |
+
thinking-token savings persist at every level.
|
| 167 |
+
|
| 168 |
+
<table class="swift15-table swift15-compact">
|
| 169 |
+
<thead>
|
| 170 |
+
<tr>
|
| 171 |
+
<th class="benchmark-heading" style="width:34%;text-align:left;padding-left:18px;">Reasoning effort</th>
|
| 172 |
+
<th class="score-heading" style="width:22%;">Qwen3.8</th>
|
| 173 |
+
<th class="score-heading" style="width:22%;">Swift 1.5</th>
|
| 174 |
+
<th class="tokens-heading" style="width:22%;">Mean thinking reduction</th>
|
| 175 |
+
</tr>
|
| 176 |
+
</thead>
|
| 177 |
+
<tbody>
|
| 178 |
+
<tr><td class="model">Xhigh</td><td>88.28%</td><td class="swift15"><strong>88.59%</strong></td><td class="change">↓ 41.9%</td></tr>
|
| 179 |
+
<tr><td class="model">Medium</td><td><strong>84.14%</strong></td><td class="swift15">82.22%</td><td class="change">↓ 24.8%</td></tr>
|
| 180 |
+
<tr><td class="model">Low</td><td>84.04%</td><td class="swift15"><strong>84.85%</strong></td><td class="change">↓ 28.7%</td></tr>
|
| 181 |
+
</tbody>
|
| 182 |
+
</table>
|
| 183 |
+
|
| 184 |
+
At `low`, Swift 1.5 scores above the base while using about 29% fewer thinking tokens.
|
| 185 |
+
|
| 186 |
+
## Quantized Swift 1.5 models
|
| 187 |
+
|
| 188 |
+
| Format | Repository | Runtime |
|
| 189 |
+
| --- | --- | --- |
|
| 190 |
+
| GGUF | [Swift-1.5-Qwen3.8-27B-GGUF](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GGUF) | llama.cpp |
|
| 191 |
+
| GSQ-RCO GGUF (compact 2–3 bit) | [Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF) | llama.cpp |
|
| 192 |
+
| AWQ INT4 (W4A16) | [Swift-1.5-Qwen3.8-27b-W4A16-AWQ](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ) | vLLM (`compressed-tensors`) |
|
| 193 |
+
| AutoRound INT4 (W4A16) | [Swift-1.5-Qwen3.8-27b-W4A16-AutoRound](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound) | vLLM (`auto-round`) |
|
| 194 |
+
| AWQ + GPTQ INT4 (W4A16) | [Swift-1.5-Qwen3.8-27b-INT4](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4) | vLLM (`compressed-tensors`) |
|
| 195 |
+
| NVFP4 | [Swift-1.5-Qwen3.8-27b-NVFP4](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-NVFP4) | NVIDIA Blackwell |
|
| 196 |
+
| AMD Quark FP8 (W8A8) | [Swift-1.5-Qwen3.8-27b-Quark-FP8-dynamic-AMD](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-Quark-FP8-dynamic-AMD) | AMD Quark |
|
| 197 |
+
| MLX 5-bit | [Swift-1.5-5bit-MLX](https://huggingface.co/ukisai/Swift-1.5-5bit-MLX) | Apple MLX |
|
| 198 |
+
| MLX 4-bit | [Swift-1.5-4bit-MLX](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX) | Apple MLX |
|
| 199 |
+
| MLX 3-bit (text only) | [Swift-1.5-3bit-MLX-TextOnly](https://huggingface.co/ukisai/Swift-1.5-3bit-MLX-TextOnly) | Apple MLX |
|
| 200 |
+
|
| 201 |
+
These results evaluate the merged Swift 1.5 checkpoint and three INT4 exports on
|
| 202 |
+
**GPQA-Diamond (198 questions), IFBench (300 prompts), and AIME 2026 (30 problems)**.
|
| 203 |
+
Each model completed the full datasets with **one sample per prompt, seed 0, and
|
| 204 |
+
zero request errors**. This is a single-seed evaluation, separate from the
|
| 205 |
+
five-repeat BF16 release results above.
|
| 206 |
+
|
| 207 |
+
The Qwen-base columns use the **saved seed/sample 0 runs**.
|
| 208 |
+
Quantization recipes and serving settings differ from the new Swift 1.5 runs,
|
| 209 |
+
so these are reference comparisons rather than a controlled measurement of
|
| 210 |
+
the Swift adaptation. Token reductions below are recomputed from those same
|
| 211 |
+
reference samples.
|
| 212 |
+
|
| 213 |
+
<table class="swift15-table" style="display:table;width:100%;table-layout:fixed;">
|
| 214 |
+
<thead><tr>
|
| 215 |
+
<th class="benchmark-heading" style="width:32%;text-align:left;padding-left:18px;">Benchmark / Swift 1.5 quantization</th>
|
| 216 |
+
<th class="score-heading" style="width:16%;">Qwen base<br>accuracy</th>
|
| 217 |
+
<th class="score-heading" style="width:16%;">Swift 1.5 quant<br>accuracy</th>
|
| 218 |
+
<th class="tokens-heading" style="width:18%;">Mean token reduction</th>
|
| 219 |
+
<th class="median-heading" style="width:18%;">Median token reduction</th>
|
| 220 |
+
</tr></thead>
|
| 221 |
+
<tbody>
|
| 222 |
+
<tr><td class="model">GPQA-Diamond<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ">AWQ INT4</a></span></td><td>86.36%</td><td class="swift15">88.38%</td><td class="change">↓ 51.5%</td><td class="change">↓ 64.4%</td></tr>
|
| 223 |
+
<tr><td class="model">GPQA-Diamond<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound">AutoRound INT4</a></span></td><td>86.36%</td><td class="swift15">89.39%</td><td class="change">↓ 50.5%</td><td class="change">↓ 57.8%</td></tr>
|
| 224 |
+
<tr><td class="model">GPQA-Diamond<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4">AWQ + GPTQ INT4</a></span></td><td>86.36%</td><td class="swift15">90.91%</td><td class="change">↓ 45.8%</td><td class="change">↓ 64.4%</td></tr>
|
| 225 |
+
<tr><td class="model">IFBench<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ">AWQ INT4</a></span></td><td>72.00%</td><td class="swift15">72.00%</td><td class="change">↓ 36.9%</td><td class="change">↓ 49.3%</td></tr>
|
| 226 |
+
<tr><td class="model">IFBench<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound">AutoRound INT4</a></span></td><td>72.00%</td><td class="swift15">69.33%</td><td class="change">↓ 29.3%</td><td class="change">↓ 39.2%</td></tr>
|
| 227 |
+
<tr><td class="model">IFBench<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4">AWQ + GPTQ INT4</a></span></td><td>72.00%</td><td class="swift15">70.00%</td><td class="change">↓ 31.8%</td><td class="change">↓ 52.7%</td></tr>
|
| 228 |
+
<tr><td class="model">AIME 2026<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ">AWQ INT4</a></span></td><td>70.00%</td><td class="swift15">86.67%</td><td class="change">↓ 29.2%</td><td class="change">↓ 36.2%</td></tr>
|
| 229 |
+
<tr><td class="model">AIME 2026<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound">AutoRound INT4</a></span></td><td>76.67%</td><td class="swift15">83.33%</td><td class="change">↓ 17.7%</td><td class="change">↓ 32.4%</td></tr>
|
| 230 |
+
<tr><td class="model">AIME 2026<br><span class="detail"><a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b-INT4">AWQ + GPTQ INT4</a></span></td><td>76.67%</td><td class="swift15">83.33%</td><td class="change">↓ 22.4%</td><td class="change">↓ 34.0%</td></tr>
|
| 231 |
+
</tbody>
|
| 232 |
+
</table>
|
| 233 |
+
|
| 234 |
+
**AIME scoring:** truncated responses count as incorrect for both columns.
|
| 235 |
+
|
| 236 |
+
The AMD Quark INT4 and FP8 exports have separate sanity evaluations; completed
|
| 237 |
+
results on these three reasoning benchmarks are not available for them.
|
| 238 |
+
|
| 239 |
+
<details>
|
| 240 |
+
<summary><strong>Quantized evaluation settings and BF16 reference</strong></summary>
|
| 241 |
+
|
| 242 |
+
**Serving:** vLLM 0.29.0, tensor parallelism 1, eager execution, BF16 activations,
|
| 243 |
+
context 131,072, template-default thinking without an effort override. The
|
| 244 |
+
AWQ + GPTQ export uses FP8 KV cache; BF16, AWQ, and AutoRound use auto KV dtype.
|
| 245 |
+
**Sampling:** temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0,
|
| 246 |
+
repetition penalty 1, seed 0. Output caps: GPQA 100,000, IFBench 81,920,
|
| 247 |
+
AIME 32,768. IFBench uses official strict prompt-level scoring.
|
| 248 |
+
|
| 249 |
+
GPQA token counts cover re-tokenized reasoning; IFBench and AIME count the
|
| 250 |
+
full generated response. Statistics include all responses, including truncations; medians
|
| 251 |
+
use the midpoint of the two central values when the sample count is even.
|
| 252 |
+
|
| 253 |
+
Saved Qwen references: W4A16 for GPQA and IFBench; Qwen AWQ for the
|
| 254 |
+
AWQ AIME row; Qwen W4A16 for the AutoRound and AWQ + GPTQ AIME rows. The
|
| 255 |
+
latter is a W4A16 reference for AutoRound, not an AutoRound base run.
|
| 256 |
+
The new runs do not reproduce the original software stack.
|
| 257 |
+
|
| 258 |
+
The fresh Swift 1.5 BF16 reference and all quantized exports scored as follows
|
| 259 |
+
under this single-seed protocol:
|
| 260 |
+
|
| 261 |
+
| Model | GPQA-Diamond | IFBench strict | AIME 2026 |
|
| 262 |
+
| --- | ---: | ---: | ---: |
|
| 263 |
+
| Swift 1.5 BF16 | 91.41% | 72.00% | 86.67% |
|
| 264 |
+
| AWQ INT4 | 88.38% | 72.00% | 86.67% |
|
| 265 |
+
| AutoRound INT4 | 89.39% | 69.33% | 83.33% |
|
| 266 |
+
| AWQ + GPTQ INT4 | 90.91% | 70.00% | 83.33% |
|
| 267 |
+
|
| 268 |
+
Truncation counts are recorded in the linked evaluation data.
|
| 269 |
+
These single-seed results do not establish
|
| 270 |
+
quality parity or replace the broader multi-seed evaluation.
|
| 271 |
+
|
| 272 |
+
[Verified counts, token statistics, settings, and evidence hashes](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/benchmarks/quantization-20260921.json).
|
| 273 |
+
|
| 274 |
+
</details>
|
| 275 |
+
|
| 276 |
+
## How to use
|
| 277 |
+
|
| 278 |
+
### GGUF download
|
| 279 |
+
|
| 280 |
+
The **[GGUF version](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GGUF)** is
|
| 281 |
+
available for compatible llama.cpp-based runtimes. For the smallest files, use the
|
| 282 |
+
**[GSQ-RCO GGUF version](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF)**: 8–12 GB mixed-precision quants refined for Swift 1.5.
|
| 283 |
+
|
| 284 |
+
### UkisAI API
|
| 285 |
+
|
| 286 |
+
Swift is served through an OpenAI-compatible API at
|
| 287 |
+
`https://ukisai.com/api/swift/v1`. It is **free for research purposes** and needs no
|
| 288 |
+
API key. The model id is `swift`.
|
| 289 |
+
|
| 290 |
+
~~~python
|
| 291 |
+
from openai import OpenAI
|
| 292 |
+
|
| 293 |
+
client = OpenAI(base_url="https://ukisai.com/api/swift/v1", api_key="none")
|
| 294 |
+
response = client.chat.completions.create(
|
| 295 |
+
model="swift",
|
| 296 |
+
messages=[{"role": "user", "content": "Explain speculative decoding in two sentences."}],
|
| 297 |
+
)
|
| 298 |
+
print(response.choices[0].message.content)
|
| 299 |
+
~~~
|
| 300 |
+
|
| 301 |
+
~~~bash
|
| 302 |
+
curl https://ukisai.com/api/swift/v1/chat/completions \
|
| 303 |
+
-H "Content-Type: application/json" \
|
| 304 |
+
-d '{"model": "swift", "messages": [{"role": "user", "content": "Hello, Swift."}]}'
|
| 305 |
+
~~~
|
| 306 |
+
|
| 307 |
+
### Transformers
|
| 308 |
+
|
| 309 |
+
~~~python
|
| 310 |
+
import torch
|
| 311 |
+
from transformers import AutoModelForImageTextToText, AutoProcessor
|
| 312 |
+
|
| 313 |
+
model_id = "ukisai/Swift-1.5-Qwen3.8-27b"
|
| 314 |
+
processor = AutoProcessor.from_pretrained(model_id)
|
| 315 |
+
model = AutoModelForImageTextToText.from_pretrained(
|
| 316 |
+
model_id,
|
| 317 |
+
torch_dtype=torch.bfloat16,
|
| 318 |
+
device_map="auto",
|
| 319 |
+
)
|
| 320 |
+
~~~
|
| 321 |
+
|
| 322 |
+
### vLLM
|
| 323 |
+
|
| 324 |
+
~~~bash
|
| 325 |
+
vllm serve ukisai/Swift-1.5-Qwen3.8-27b \
|
| 326 |
+
--dtype bfloat16 \
|
| 327 |
+
--tensor-parallel-size 1 \
|
| 328 |
+
--max-model-len 262144 \
|
| 329 |
+
--reasoning-parser qwen3 \
|
| 330 |
+
--enable-auto-tool-choice \
|
| 331 |
+
--tool-call-parser qwen3_coder \
|
| 332 |
+
--port 8000
|
| 333 |
+
~~~
|
| 334 |
+
|
| 335 |
+
### SGLang
|
| 336 |
+
|
| 337 |
+
Alternatively, use a current SGLang build with Qwen3.8 support:
|
| 338 |
+
|
| 339 |
+
~~~bash
|
| 340 |
+
python -m sglang.launch_server \
|
| 341 |
+
--model-path ukisai/Swift-1.5-Qwen3.8-27b \
|
| 342 |
+
--dtype bfloat16 \
|
| 343 |
+
--tp-size 1 \
|
| 344 |
+
--context-length 262144 \
|
| 345 |
+
--reasoning-parser qwen3 \
|
| 346 |
+
--tool-call-parser qwen3_coder \
|
| 347 |
+
--port 8000
|
| 348 |
+
~~~
|
| 349 |
+
|
| 350 |
+
Adjust tensor parallelism and context length to your GPU memory. See the base model's
|
| 351 |
+
[vLLM recipe](https://recipes.vllm.ai/Qwen/Qwen3.8-27B) and
|
| 352 |
+
[SGLang recipe](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B)
|
| 353 |
+
for installation and hardware-specific settings.
|
| 354 |
+
|
| 355 |
+
### Optional MTP decoding
|
| 356 |
+
|
| 357 |
+
The published weights include the base model's MTP head. To enable self-speculative
|
| 358 |
+
decoding, append the corresponding flags to the server command above:
|
| 359 |
+
|
| 360 |
+
~~~bash
|
| 361 |
+
# vLLM
|
| 362 |
+
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
|
| 363 |
+
|
| 364 |
+
# SGLang
|
| 365 |
+
--speculative-algorithm EAGLE --speculative-num-steps 3 \
|
| 366 |
+
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4
|
| 367 |
+
~~~
|
| 368 |
+
|
| 369 |
+
## License and access
|
| 370 |
+
|
| 371 |
+
Swift 1.5 is a derivative of [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)
|
| 372 |
+
(Copyright 2026 Alibaba Cloud, [Apache License 2.0](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE-APACHE-2.0)). UkisAI's contribution, including the adapted
|
| 373 |
+
weights, is licensed under the **[Swift Open License v1.0](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE)**. See [NOTICE](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/NOTICE) for the change notice and attribution details.
|
| 374 |
+
|
| 375 |
+
Personal, research, educational, evaluation, and commercial use are free for individuals
|
| 376 |
+
and organizations with gross annual revenue, including affiliates, of up to US$1,000,000.
|
| 377 |
+
Above that threshold, commercial use requires a separate Swift Enterprise License.
|
| 378 |
+
Contact [UkisAI](https://ukisai.com/contact) for terms.
|
| 379 |
+
|
| 380 |
+
Nothing in the Swift Open License limits rights in Qwen3.8-27B itself under Apache 2.0.
|
| 381 |
+
|
| 382 |
+
## Citation
|
| 383 |
+
|
| 384 |
+
~~~bibtex
|
| 385 |
+
@misc{swift-1.5-qwen3.8-27b,
|
| 386 |
+
title = {Swift 1.5 Qwen3.8-27B},
|
| 387 |
+
author = {UkisAI},
|
| 388 |
+
year = {2026},
|
| 389 |
+
url = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
|
| 390 |
+
}
|
| 391 |
+
~~~
|
| 392 |
+
|
| 393 |
+
## Acknowledgements
|
| 394 |
+
|
| 395 |
+
We acknowledge the [NVIDIA Innovation Lab](https://www.nvidia.com/en-us/data-center/innovation-lab/),
|
| 396 |
+
[Amazon Web Services](https://aws.amazon.com/), and
|
| 397 |
+
[Google Cloud](https://cloud.google.com/) for providing the
|
| 398 |
+
compute for Swift's development, training, and evaluation.
|