Swift-1.5-Qwen3.8-27B-PARO-MXFP6

ParoQuant MXFP6 quant of ukisai/Swift-1.5-Qwen3.8-27b. Unofficial.

  • Scheme: MXFP6 E2M3 weights (W6A8), ParoQuant rotations (krot 8, group 128), quant_method: paroquant_mxfp6
  • Method: rotations from z-lab/Qwen3.8-27B-PARO (trained on base Qwen3.8-27B, frozen), then round-to-nearest to MXFP6. Stage-2 fine-tune skipped for time.
  • Same recipe as hugypufy/Swift-Qwen3.8-27B-PARO-MXFP6, minus its fine-tune
  • paroquant_mxfp6 is only supported by some custom RDNA4 vLLM builds for now; stock vLLM will not load it
  • ~24 GB

Original model card

UkisAI
Website  •  Learn more  •  GGUF  •  GSQ-RCO GGUF  •  Evaluation  •  Enterprise licensing

Swift 1.5 Qwen3.8-27B

Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B. It uses 58.5% fewer thinking tokens while scoring 0.35% higher than the base, for a 1.95× speed-up on several tasks.

Swift 1.5 is a direct upgrade from Swift 1.0, our model with 350k+ downloads, delivering stronger overall performance than both base and Swift 1.0 in various tasks, especially coding and agentic, while using fewer thinking tokens. We accomplished that by scaling up the post-training (RL and OPD) from the previous version.

Demo

We gave base Qwen3.8-27B and Swift 1.5 27B the same prompt:

create a 3d little planet globe where I (player can walk around) and it has all these biomes to explore, the globe doesn't have to be too big, but still fun to go around. It's about a boy scout who is camping and goes around exploring.

Try the game yourself here: https://ukisai.com/swift-games/27b

Base Qwen3.8-27B took 104.6 minutes to build its game. Swift 1.5 took 11.39 minutes.

Training approach

We made Swift efficient by figuring out which tokens were linked to pathological overthinking and penalizing them without "attacking" the reasoning length directly then regained the accuracy with RL and OPD, leading to "compressed" token usage while maintaining accuracy. Swift 1.5 was made from Swift 1.0, on whom we scaled up the post-training methods that previously improved Swift1.0 model performance, this time with the main focus on long-horizon, agentic, and coding tasks, as seen in the LiveCodeBench and Terminal Bench 2.1 improvements. Our training data is viewable here: https://huggingface.co/datasets/ukisai/Qwen3.8-27B-multi-turn-agent-sft albeit it is not used out of the box, but rather re-sampled, turned into proper RL environments etc.

Evaluation

The external results below compare Qwen3.8-27B, the foundation base model, and Swift 1.5. Both models use the same saved evaluation protocols, and all scores are reported as final aggregate percentages.

Benchmark Final score Mean tokens Median tokens
Qwen3.8 Swift 1.5 Qwen3.8 Swift 1.5 Reduction Reduction
General reasoning
GPQA-Diamond88.28%88.59%15,0148,717↓ 41.9%↓ 58.5%
C-Eval90.00%90.92%1,492819↓ 45.1%↓ 16.9%
IFBench73.53%72.07%8,0524,955↓ 38.5%↓ 47.3%
ERQA67.45%65.40%4,1371,906↓ 53.9%↓ 56.2%
Mathematics
AIME 202698.67%96.00%22,01413,203↓ 40.0%↓ 48.5%
HMMT November 202599.33%97.33%22,03214,957↓ 32.1%↓ 47.8%
Coding
LiveCodeBench v676.76%81.71%11,1848,448↓ 24.5%↓ 46.3%
Agent tasks
Terminal-Bench 2.1*69.21%72.13%52,26543,733↓ 16.3%↓ 0.1%

* Note: Terminal Bench 2.1 score of Swift1.5 27B is misleadingly low at first glance. It is not a bug, but a simple matter of the Swift models not falling into overthinking loops and failing the task, rather pursuing it until the end, leading to higher average token usage. The token reduction still falls in the -38.7% range when compared apples-to-apples.

Benchmark methodology and reproduction settings

Serving: BF16 · vLLM 0.27.1 · Qwen3 parser · context 262,144 · thinking xhigh.
Sampling: temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0 · presence_penalty 0 · repetition_penalty 1.
Benchmarks: averages over five seeds (0–4) per model; five trials per task for Terminal-Bench, base and Swift 1.5 served at context 131,072 on the same Harbor build.

BenchmarkOutput cap
GPQA-Diamond100,000
C-Eval16,384
IFBench81,920
ERQA100,000
AIME 2026250,000
HMMT November 2025250,000
LiveCodeBench v632,768
Terminal-Bench 2.1Agent/task limits

Efficiency across reasoning efforts

Qwen3.8's reasoning_effort setting lets users choose how much the model thinks. For Swift 1.5 to be useful across these settings, it needs to reduce thinking while keeping accuracy close to the base. We therefore tested xhigh, medium, and low: thinking-token savings persist at every level.

Reasoning effort Qwen3.8 Swift 1.5 Mean thinking reduction
Xhigh88.28%88.59%↓ 41.9%
Medium84.14%82.22%↓ 24.8%
Low84.04%84.85%↓ 28.7%

At low, Swift 1.5 scores above the base while using about 29% fewer thinking tokens.

Quantized Swift 1.5 models

Format Repository Runtime
GGUF Swift-1.5-Qwen3.8-27B-GGUF llama.cpp
GSQ-RCO GGUF (compact 2–3 bit) Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF llama.cpp
AWQ INT4 (W4A16) Swift-1.5-Qwen3.8-27b-W4A16-AWQ vLLM (compressed-tensors)
AutoRound INT4 (W4A16) Swift-1.5-Qwen3.8-27b-W4A16-AutoRound vLLM (auto-round)
AWQ + GPTQ INT4 (W4A16) Swift-1.5-Qwen3.8-27b-INT4 vLLM (compressed-tensors)
NVFP4 Swift-1.5-Qwen3.8-27b-NVFP4 NVIDIA Blackwell
AMD Quark FP8 (W8A8) Swift-1.5-Qwen3.8-27b-Quark-FP8-dynamic-AMD AMD Quark
MLX 5-bit Swift-1.5-5bit-MLX Apple MLX
MLX 4-bit Swift-1.5-4bit-MLX Apple MLX
MLX 3-bit (text only) Swift-1.5-3bit-MLX-TextOnly Apple MLX

These results evaluate the merged Swift 1.5 checkpoint and three INT4 exports on GPQA-Diamond (198 questions), IFBench (300 prompts), and AIME 2026 (30 problems). Each model completed the full datasets with one sample per prompt, seed 0, and zero request errors. This is a single-seed evaluation, separate from the five-repeat BF16 release results above.

The Qwen-base columns use the saved seed/sample 0 runs. Quantization recipes and serving settings differ from the new Swift 1.5 runs, so these are reference comparisons rather than a controlled measurement of the Swift adaptation. Token reductions below are recomputed from those same reference samples.

Benchmark / Swift 1.5 quantization Qwen base
accuracy
Swift 1.5 quant
accuracy
Mean token reduction Median token reduction
GPQA-Diamond
AWQ INT4
86.36%88.38%↓ 51.5%↓ 64.4%
GPQA-Diamond
AutoRound INT4
86.36%89.39%↓ 50.5%↓ 57.8%
GPQA-Diamond
AWQ + GPTQ INT4
86.36%90.91%↓ 45.8%↓ 64.4%
IFBench
AWQ INT4
72.00%72.00%↓ 36.9%↓ 49.3%
IFBench
AutoRound INT4
72.00%69.33%↓ 29.3%↓ 39.2%
IFBench
AWQ + GPTQ INT4
72.00%70.00%↓ 31.8%↓ 52.7%
AIME 2026
AWQ INT4
70.00%86.67%↓ 29.2%↓ 36.2%
AIME 2026
AutoRound INT4
76.67%83.33%↓ 17.7%↓ 32.4%
AIME 2026
AWQ + GPTQ INT4
76.67%83.33%↓ 22.4%↓ 34.0%

AIME scoring: truncated responses count as incorrect for both columns.

The AMD Quark INT4 and FP8 exports have separate sanity evaluations; completed results on these three reasoning benchmarks are not available for them.

Quantized evaluation settings and BF16 reference

Serving: vLLM 0.29.0, tensor parallelism 1, eager execution, BF16 activations, context 131,072, template-default thinking without an effort override. The AWQ + GPTQ export uses FP8 KV cache; BF16, AWQ, and AutoRound use auto KV dtype. Sampling: temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0, repetition penalty 1, seed 0. Output caps: GPQA 100,000, IFBench 81,920, AIME 32,768. IFBench uses official strict prompt-level scoring.

GPQA token counts cover re-tokenized reasoning; IFBench and AIME count the full generated response. Statistics include all responses, including truncations; medians use the midpoint of the two central values when the sample count is even.

Saved Qwen references: W4A16 for GPQA and IFBench; Qwen AWQ for the AWQ AIME row; Qwen W4A16 for the AutoRound and AWQ + GPTQ AIME rows. The latter is a W4A16 reference for AutoRound, not an AutoRound base run. The new runs do not reproduce the original software stack.

The fresh Swift 1.5 BF16 reference and all quantized exports scored as follows under this single-seed protocol:

Model GPQA-Diamond IFBench strict AIME 2026
Swift 1.5 BF16 91.41% 72.00% 86.67%
AWQ INT4 88.38% 72.00% 86.67%
AutoRound INT4 89.39% 69.33% 83.33%
AWQ + GPTQ INT4 90.91% 70.00% 83.33%

Truncation counts are recorded in the linked evaluation data. These single-seed results do not establish quality parity or replace the broader multi-seed evaluation.

Verified counts, token statistics, settings, and evidence hashes.

How to use

GGUF download

The GGUF version is available for compatible llama.cpp-based runtimes. For the smallest files, use the GSQ-RCO GGUF version: 8–12 GB mixed-precision quants refined for Swift 1.5.

UkisAI API

Swift is served through an OpenAI-compatible API at https://ukisai.com/api/swift/v1. It is free for research purposes and needs no API key. The model id is swift.

from openai import OpenAI

client = OpenAI(base_url="https://ukisai.com/api/swift/v1", api_key="none")
response = client.chat.completions.create(
    model="swift",
    messages=[{"role": "user", "content": "Explain speculative decoding in two sentences."}],
)
print(response.choices[0].message.content)
curl https://ukisai.com/api/swift/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "swift", "messages": [{"role": "user", "content": "Hello, Swift."}]}'

Transformers

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "ukisai/Swift-1.5-Qwen3.8-27b"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

vLLM

vllm serve ukisai/Swift-1.5-Qwen3.8-27b \
  --dtype bfloat16 \
  --tensor-parallel-size 1 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --port 8000

SGLang

Alternatively, use a current SGLang build with Qwen3.8 support:

python -m sglang.launch_server \
  --model-path ukisai/Swift-1.5-Qwen3.8-27b \
  --dtype bfloat16 \
  --tp-size 1 \
  --context-length 262144 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --port 8000

Adjust tensor parallelism and context length to your GPU memory. See the base model's vLLM recipe and SGLang recipe for installation and hardware-specific settings.

Optional MTP decoding

The published weights include the base model's MTP head. To enable self-speculative decoding, append the corresponding flags to the server command above:

# vLLM
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'

# SGLang
--speculative-algorithm EAGLE --speculative-num-steps 3 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 4

License and access

Swift 1.5 is a derivative of Qwen3.8-27B (Copyright 2026 Alibaba Cloud, Apache License 2.0). UkisAI's contribution, including the adapted weights, is licensed under the Swift Open License v1.0. See NOTICE for the change notice and attribution details.

Personal, research, educational, evaluation, and commercial use are free for individuals and organizations with gross annual revenue, including affiliates, of up to US$1,000,000. Above that threshold, commercial use requires a separate Swift Enterprise License. Contact UkisAI for terms.

Nothing in the Swift Open License limits rights in Qwen3.8-27B itself under Apache 2.0.

Citation

@misc{swift-1.5-qwen3.8-27b,
  title  = {Swift 1.5 Qwen3.8-27B},
  author = {UkisAI},
  year   = {2026},
  url    = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}

Acknowledgements

We acknowledge the NVIDIA Innovation Lab, Amazon Web Services, and Google Cloud for providing the compute for Swift's development, training, and evaluation.

Downloads last month
9
Safetensors
Model size
21B params
Tensor type
F16
·
I16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6

Base model

Qwen/Qwen3.8-27B
Quantized
(29)
this model

Paper for realderpz/Swift-1.5-Qwen3.8-27B-PARO-MXFP6