Instructions to use rovangju/Swift-Qwen3.8-27b-W8A8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rovangju/Swift-Qwen3.8-27b-W8A8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="rovangju/Swift-Qwen3.8-27b-W8A8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("rovangju/Swift-Qwen3.8-27b-W8A8") model = AutoModelForMultimodalLM.from_pretrained("rovangju/Swift-Qwen3.8-27b-W8A8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use rovangju/Swift-Qwen3.8-27b-W8A8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "rovangju/Swift-Qwen3.8-27b-W8A8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rovangju/Swift-Qwen3.8-27b-W8A8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/rovangju/Swift-Qwen3.8-27b-W8A8
- SGLang
How to use rovangju/Swift-Qwen3.8-27b-W8A8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "rovangju/Swift-Qwen3.8-27b-W8A8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rovangju/Swift-Qwen3.8-27b-W8A8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "rovangju/Swift-Qwen3.8-27b-W8A8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rovangju/Swift-Qwen3.8-27b-W8A8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use rovangju/Swift-Qwen3.8-27b-W8A8 with Docker Model Runner:
docker model run hf.co/rovangju/Swift-Qwen3.8-27b-W8A8
Swift-Qwen3.8-27b-W8A8
W8A8 (INT8 weights and INT8 activations) post-training quantization of
ukisai/Swift-Qwen3.8-27b — a
27B multimodal (text + vision) model in the Qwen3.5 family
(Qwen3_5ForConditionalGeneration).
Benchmarks
Head-to-head against the SmoothQuant W8A8 checkpoint (Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8) — our running "champ", i.e. the best-performing W8A8 configuration for this model in our benchmark harness — with an identical serve config (FLASH_ATTN, DFlash K=5 speculative decoding, no KV-cache dtype change), concurrency ≤ 3, full runs:
| Metric | Swift W8A8 (this repo) | SmoothQuant W8A8 | Δ |
|---|---|---|---|
| Prefill, 8192→64 — TTFT (median) | 2175.4 ms | 2583.1 ms | −15.8% |
| Prefill, 8192→64 — tok/s | 3766 | 3171 | +18.8% |
| Decode, 256→1024 — TPOT (median) | 9.98 ms | 9.34 ms | +6.9% |
| Decode, 256→1024 — tok/s | 100.2 | 107.1 | −6.5% |
| Balanced — TTFT (median) | 647.9 ms | 794.0 ms | −18.4% |
| Balanced — TPOT (median) | 14.97 ms | 16.71 ms | −10.4% |
| Balanced — output tok/s | 152.8 | 141.8 | +7.7% |
Benchmarked on a single NVIDIA CMP 170HX (64 GB) — the study used GPU 1 of a 2-GPU box, Xeon Gold 6154 host.
Headline: Swift W8A8 wins the operating point that matters (balanced agent workload: +7.7% output tok/s, −18.4% TTFT) and prefill by a wide margin, at the cost of ~6.5% raw decode tok/s. Note the benchmark harness holds output lengths fixed, so it does not capture Swift's main real-world advantage — shorter replies / less thinking per turn.
Quantization details
| Format | compressed-tensors, loadable by vLLM via CompressedTensorsW8A8Int8 |
| Weights | INT8, symmetric, per-channel, static scales |
| Activations | INT8, symmetric, per-token, dynamic at serve time (no stored input scales, no calibration data) |
| GEMMs | True INT8×INT8 at inference (activations are not upcast to bf16, unlike W8A16) |
| Kept in bf16 | vision tower (model.visual.*), lm_head (~1.3B params), MTP module |
| Quantized with | llm-compressor 0.14 oneshot, QuantizationModifier(targets="Linear", scheme="W8A8") |
| Calibration | None required — weight scales derive from the weights themselves and activation quantization is dynamic |
| Checkpoint size | ~29.1 GiB |
Note: this is not a SmoothQuant checkpoint. It is per-channel weights + dynamic activations. It is the same compressed-tensors family as static SmoothQuant W8A8 checkpoints and loads identically in vLLM.
Two things were handled deliberately during quantization:
- Full multimodal load.
llm-compressor's defaultoneshot(model="...")path loads viaAutoModelForCausalLM, which resolvesqwen3_5to the text-only class and silently drops the vision tower (333 tensors). The fullAutoModelForImageTextToTextmodel was loaded and handed tooneshot()directly, so allmodel.visual.*tensors are present (and kept in bf16). - Post-save verification (tensor-for-tensor accounting against the original index, per-channel scale shapes checked, and a hard failure if any activation scale tensors were written) passed for all shards.
Usage — vLLM
vllm serve rovangju/Swift-Qwen3.8-27b-W8A8 --dtype bfloat16 --enable-prefix-caching
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="rovangju/Swift-Qwen3.8-27b-W8A8",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
(Tensor-parallel across two GPUs also works; single-GPU TP=1 is the default.)
Quality checks
Run with lm-eval-harness against the vLLM backend
(dtype=bfloat16, max_model_len=4096, seed 1234, 0-shot), on a
200-sample smoke subset of each task:
| Task | Metric | Value | Stderr |
|---|---|---|---|
| ARC-Easy | acc | 0.795 | ± 0.0286 |
| ARC-Easy | acc_norm | 0.680 | ± 0.0331 |
| HellaSwag | acc | 0.605 | ± 0.0347 |
| HellaSwag | acc_norm | 0.740 | ± 0.0311 |
| TruthfulQA (MC2) | acc | 0.5394 | ± 0.0308 |
| GSM8K (3-shot) | exact_match (flexible-extract) | 0.565 | ± 0.0351 |
| GSM8K (3-shot) | exact_match (strict-match) | 0.000 | ± 0.0000 |
These are smoke-check numbers from a 200-sample subset — the stderr margins are wide; treat them as a sanity gate, not a final benchmark. The strict-match GSM8K score is a known artifact of the harness's strict filter on this template, not a model failure.
Appendix — benchmark serve script
The vLLM run used for the benchmark numbers above (identical for both models; the only difference between the two runs was the model path):
#!/usr/bin/env bash
set -euo pipefail
vllm serve <MODEL> \
--dtype bfloat16 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--max-model-len auto \
--max-num-batched-tokens 16384 \
--attention-config '{"backend": "FLASH_ATTN"}' \
--mamba-cache-mode align \
--enable-prefix-caching \
--per-request-spec-decode-metrics summary \
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":5,"draft_sample_method":"probabilistic"}'
<MODEL> is Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8 for the
comparison column and this repo (models/Swift-Qwen3.8-27b-W8A8) for the
benchmark column. The draft model in the speculative config is the DFlash
draft for Qwen3.8-27B; note the DFlash draft was not re-finetuned for the
Swift distribution.
License
swift-open-license-1.0, inherited from the base model
ukisai/Swift-Qwen3.8-27b —
see the base model's LICENSE
for the exact terms (this quantization adds no restrictions of its own).
- Downloads last month
- 11