--- license: other license_name: swift-open-license-1.0 license_link: https://huggingface.co/ukisai/Swift-Qwen3.8-27b/blob/main/LICENSE library_name: transformers pipeline_tag: image-text-to-text gated: true tags: - qwen3_8 - efficient-thinking - reasoning - token-efficient - lora base_model: ukisai/Swift-Qwen3.8-27b base_model_relation: quantized ---
UkisAI
Website  •  Learn more  •  GGUF  •  Enterprise licensing
# Swift-Qwen3.8-27B Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, using **58.3% fewer thinking tokens** while maintaining near-identical performance (**<1% loss**) and as a result getting a **x1.95 speed-up** on several tasks.

The prompt is a sample from LiveCodeBench v6

## Training approach We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in Qwen’s reasoning rollouts. We then fine-tuned Qwen by penalizing usage of those tokens while it reasons. Swift produces shorter reasoning traces. In our testing, we also observe fewer overthinking errors. For maximum gains, Swift also includes a transfer component derived from [BottleCap AI's ThinkingCap-Qwen3.6-27B](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B). ## Evaluation scope > All results below compare the Qwen3.8-27B BF16 base with the same base plus the > Swift adapter. ## Benchmarks
Benchmark Score Mean tokens Median tokens
Base Swift Base Swift Reduction Reduction
General reasoning
GPQA-Diamond88.38%88.28%15,0148,855↓ 41.0%↓ 58.3%
MMLU-Pro85.47%84.95%2,9801,603↓ 46.2%↓ 28.3%
C-Eval90.00%90.62%1,492804↓ 46.1%↓ 19.3%
IFBench73.53%71.80%8,0524,657↓ 42.2%↓ 50.5%
Mathematics
AIME 202698.67%94.00%22,01416,143↓ 26.7%↓ 50.2%
HMMT (Nov 2025)99.33%96.00%22,03215,189↓ 31.1%↓ 45.9%
Multimodal
ERQA67.45%66.30%4,1372,045↓ 50.6%↓ 54.6%
Agentic coding
Terminal-Bench 2.166.74%65.84%37,08627,272↓ 26.5%↓ 38.7%
LiveCodeBench v676.76%81.55%11,3748,615↓ 24.3%↓ 45.8%
How to reproduce

Serving: BF16 · vLLM 0.27.1 · Qwen3 parser · context 262,144 · thinking xhigh.
Sampling: temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0 · presence_penalty 0 · repetition_penalty 1.
Benchmarks: averages over five seeds (0–4) per model; five trials per task for Terminal-Bench.

BenchmarkOutput cap
GPQA-Diamond100,000
MMLU-Pro100,000
C-Eval16,384
IFBench81,920
AIME 2026250,000
HMMT Nov 2025250,000
ERQA100,000
Terminal-Bench 2.1Agent/task limits
LiveCodeBench v632,768
## Efficiency across and versus reasoning efforts Qwen3.8's `reasoning_effort` setting lets users choose how much the model thinks. For Swift to be useful across these settings, it needs to reduce thinking while keeping accuracy close to the base. We therefore tested `xhigh`, `medium`, and `low`: thinking-token savings persist at every level.
Reasoning effort Mean thinking reduction
Xhigh↓ 41.0%
Medium↓ 22.7%
Low↓ 25.8%
The efficiency also holds up against the base's own lower effort settings. On GPQA-Diamond (198 questions, 5 seeds, 990 paired calls), Swift at `xhigh` is compared with the base at `xhigh` and at `medium`:
GPQA-Diamond Score Mean tokens Median tokens
Base · xhigh88.38%15,0146,642
Swift · xhigh88.28%8,8552,771
Base · medium84.14%4,4511,753
Swift retains the accuracy of `xhigh` while using about half the tokens, although it uses about double the tokens of `medium`. ## Quantized models Quantized deployment is the intended use for Swift: lower-memory weights paired with shorter reasoning. The INT4 evaluations below retain token savings across GPQA, IFBench, and AIME. On AIME, Swift matches or improves accuracy and reduces output-cap failures by **31–33%**.
Benchmark / quantization Base accuracy Swift accuracy Mean token reduction Median token reduction
GPQA-Diamond
Mixed-precision quant W4A16 · thinking tokens
88.69%88.38%↓ 32.1%↓ 50.2%
IFBench
Mixed-precision quant W4A16 · completion tokens
72.58%71.25%↓ 30.1%↓ 38.0%
AIME 2026
Mixed-precision quant W4A16 · completion tokens
84.00%84.00%↓ 19.0%↓ 37.5%
AIME 2026
AWQ INT4 · completion tokens
82.67%84.00%↓ 22.8%↓ 34.8%
Quantized evaluation settings Each row compares the same quantized base with and without the Swift adapter. GPQA and AIME use five seeds; IFBench uses four samples per prompt and strict scoring. Output caps: GPQA 100,000; IFBench 81,920; AIME 32,768. GPQA and IFBench use saved historical base runs. AIME uses template-default effort and counts truncated answers as incorrect. Its shorter cap makes it a separate comparison from the BF16 table.
## How to use ### GGUF download The **[GGUF version](https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF)** is available for compatible llama.cpp-based runtime. ### UkisAI API Swift is served through an OpenAI-compatible API at `https://ukisai.com/api/swift/v1`. It is **free for research purposes** and needs no API key. The model id is `swift`. ```python from openai import OpenAI client = OpenAI(base_url="https://ukisai.com/api/swift/v1", api_key="none") response = client.chat.completions.create( model="swift", messages=[{"role": "user", "content": "Explain speculative decoding in two sentences."}], ) print(response.choices[0].message.content) ``` ```bash curl https://ukisai.com/api/swift/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "swift", "messages": [{"role": "user", "content": "Hello, Swift."}]}' ``` ### Transformers ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor model_id = "ukisai/Swift-Qwen3.8-27b" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForImageTextToText.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ) ``` ### vLLM ```bash vllm serve ukisai/Swift-Qwen3.8-27b \ --dtype bfloat16 \ --tensor-parallel-size 1 \ --max-model-len 262144 \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --port 8000 ``` ### SGLang Alternatively, use a current SGLang build with Qwen3.8 support: ```bash python -m sglang.launch_server \ --model-path ukisai/Swift-Qwen3.8-27b \ --dtype bfloat16 \ --tp-size 1 \ --context-length 262144 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder \ --port 8000 ``` Adjust tensor parallelism and context length to your GPU memory. See the base model's [vLLM recipe](https://recipes.vllm.ai/Qwen/Qwen3.8-27B) and [SGLang recipe](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B) for installation and hardware-specific settings. ### Optional MTP decoding The published weights include the base model's MTP head. To enable self-speculative decoding, append the corresponding flags to the server command above: ```bash # vLLM --speculative-config '{"method":"mtp","num_speculative_tokens":3}' # SGLang --speculative-algorithm EAGLE --speculative-num-steps 3 \ --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 ``` ## License and access Swift-Qwen3.8-27B is a derivative of [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) (Copyright 2026 Alibaba Cloud, [Apache License 2.0](https://huggingface.co/ukisai/Swift-Qwen3.8-27b/blob/main/LICENSE-APACHE-2.0)). UkisAI's contribution, the fine-tuned weights, is licensed under the **[Swift Open License v1.0](https://huggingface.co/ukisai/Swift-Qwen3.8-27b/blob/main/LICENSE)**. See [NOTICE](https://huggingface.co/ukisai/Swift-Qwen3.8-27b/blob/main/NOTICE) for exactly what was changed. Personal, research, educational, evaluation, and commercial use are free for individuals and organizations with gross annual revenue, including affiliates, of up to US$1,000,000. Above that threshold, commercial use requires a separate **Swift Enterprise License**. Contact [UkisAI](https://ukisai.com/contact) for terms. Nothing in the Swift Open License limits your rights in Qwen3.8-27B itself under Apache 2.0. ## Citation ```bibtex @misc{swift-qwen3.8-27b, title = {Swift-Qwen3.8-27B}, author = {UkisAI}, year = {2026}, url = {https://huggingface.co/ukisai/Swift-Qwen3.8-27b} } ``` ## Acknowledgements We acknowledge the [NVIDIA Innovation Lab](https://www.nvidia.com/en-us/data-center/innovation-lab/) for providing access to **8× NVIDIA H100 GPUs** to train Swift.