Swift-Qwen3.8-27b-W8A8

W8A8 (INT8 weights and INT8 activations) post-training quantization of ukisai/Swift-Qwen3.8-27b — a 27B multimodal (text + vision) model in the Qwen3.5 family (Qwen3_5ForConditionalGeneration).

Benchmarks

Head-to-head against the SmoothQuant W8A8 checkpoint (Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8) — our running "champ", i.e. the best-performing W8A8 configuration for this model in our benchmark harness — with an identical serve config (FLASH_ATTN, DFlash K=5 speculative decoding, no KV-cache dtype change), concurrency ≤ 3, full runs:

Metric Swift W8A8 (this repo) SmoothQuant W8A8 Δ
Prefill, 8192→64 — TTFT (median) 2175.4 ms 2583.1 ms −15.8%
Prefill, 8192→64 — tok/s 3766 3171 +18.8%
Decode, 256→1024 — TPOT (median) 9.98 ms 9.34 ms +6.9%
Decode, 256→1024 — tok/s 100.2 107.1 −6.5%
Balanced — TTFT (median) 647.9 ms 794.0 ms −18.4%
Balanced — TPOT (median) 14.97 ms 16.71 ms −10.4%
Balanced — output tok/s 152.8 141.8 +7.7%

Benchmarked on a single NVIDIA CMP 170HX (64 GB) — the study used GPU 1 of a 2-GPU box, Xeon Gold 6154 host.

Headline: Swift W8A8 wins the operating point that matters (balanced agent workload: +7.7% output tok/s, −18.4% TTFT) and prefill by a wide margin, at the cost of ~6.5% raw decode tok/s. Note the benchmark harness holds output lengths fixed, so it does not capture Swift's main real-world advantage — shorter replies / less thinking per turn.

Quantization details

Format compressed-tensors, loadable by vLLM via CompressedTensorsW8A8Int8
Weights INT8, symmetric, per-channel, static scales
Activations INT8, symmetric, per-token, dynamic at serve time (no stored input scales, no calibration data)
GEMMs True INT8×INT8 at inference (activations are not upcast to bf16, unlike W8A16)
Kept in bf16 vision tower (model.visual.*), lm_head (~1.3B params), MTP module
Quantized with llm-compressor 0.14 oneshot, QuantizationModifier(targets="Linear", scheme="W8A8")
Calibration None required — weight scales derive from the weights themselves and activation quantization is dynamic
Checkpoint size ~29.1 GiB

Note: this is not a SmoothQuant checkpoint. It is per-channel weights + dynamic activations. It is the same compressed-tensors family as static SmoothQuant W8A8 checkpoints and loads identically in vLLM.

Two things were handled deliberately during quantization:

  • Full multimodal load. llm-compressor's default oneshot(model="...") path loads via AutoModelForCausalLM, which resolves qwen3_5 to the text-only class and silently drops the vision tower (333 tensors). The full AutoModelForImageTextToText model was loaded and handed to oneshot() directly, so all model.visual.* tensors are present (and kept in bf16).
  • Post-save verification (tensor-for-tensor accounting against the original index, per-channel scale shapes checked, and a hard failure if any activation scale tensors were written) passed for all shards.

Usage — vLLM

vllm serve rovangju/Swift-Qwen3.8-27b-W8A8 --dtype bfloat16 --enable-prefix-caching
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
    model="rovangju/Swift-Qwen3.8-27b-W8A8",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)

(Tensor-parallel across two GPUs also works; single-GPU TP=1 is the default.)

Quality checks

Run with lm-eval-harness against the vLLM backend (dtype=bfloat16, max_model_len=4096, seed 1234, 0-shot), on a 200-sample smoke subset of each task:

Task Metric Value Stderr
ARC-Easy acc 0.795 ± 0.0286
ARC-Easy acc_norm 0.680 ± 0.0331
HellaSwag acc 0.605 ± 0.0347
HellaSwag acc_norm 0.740 ± 0.0311
TruthfulQA (MC2) acc 0.5394 ± 0.0308
GSM8K (3-shot) exact_match (flexible-extract) 0.565 ± 0.0351
GSM8K (3-shot) exact_match (strict-match) 0.000 ± 0.0000

These are smoke-check numbers from a 200-sample subset — the stderr margins are wide; treat them as a sanity gate, not a final benchmark. The strict-match GSM8K score is a known artifact of the harness's strict filter on this template, not a model failure.

Appendix — benchmark serve script

The vLLM run used for the benchmark numbers above (identical for both models; the only difference between the two runs was the model path):

#!/usr/bin/env bash

set -euo pipefail

vllm serve <MODEL> \
  --dtype bfloat16 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --max-model-len auto \
  --max-num-batched-tokens 16384 \
  --attention-config '{"backend": "FLASH_ATTN"}' \
  --mamba-cache-mode align \
  --enable-prefix-caching \
  --per-request-spec-decode-metrics summary \
  --speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":5,"draft_sample_method":"probabilistic"}'

<MODEL> is Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8 for the comparison column and this repo (models/Swift-Qwen3.8-27b-W8A8) for the benchmark column. The draft model in the speculative config is the DFlash draft for Qwen3.8-27B; note the DFlash draft was not re-finetuned for the Swift distribution.

License

swift-open-license-1.0, inherited from the base model ukisai/Swift-Qwen3.8-27b — see the base model's LICENSE for the exact terms (this quantization adds no restrictions of its own).

Downloads last month
11
Safetensors
Model size
28B params
Tensor type
BF16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rovangju/Swift-Qwen3.8-27b-W8A8

Base model

Qwen/Qwen3.8-27B
Quantized
(57)
this model