Instructions to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast") model = AutoModelForMultimodalLM.from_pretrained("TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
- SGLang
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with Docker Model Runner:
docker model run hf.co/TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
Swift-Qwen3.8-27B-Uncensored-W4A16-fast
The "fast variant" of
Swift-Qwen3.8-27B-Uncensored-W4A16:
the same 4-bit body of d0xin/Swift-Qwen3.8-27B-Uncensored-BF16
(UkisAI's reasoning-efficient Swift-Qwen3.8-27B
with its refusal direction removed), with the lm_head and the MTP module requantized to
int4 GPTQ. It serves on one 24 GB GPU (RTX 3090 class) with the
HyperQwen stack. The int4 lm_head makes every
decode step read about 0.65 GB less; that is where the gain over the base build comes from.
Lineage
Qwen/Qwen3.8-27B Apache 2.0
└─ ukisai/Swift-Qwen3.8-27b Swift Open License v1.0, reasoning-efficiency LoRA merged
└─ d0xin/Swift-Qwen3.8-27B-Uncensored-BF16 rank-1 directional ablation, layer 38
└─ ...-W4A16 W4A16 g128 body, int8 lm_head / embed_tokens / MTP
└─ this model int4 GPTQ lm_head / MTP
What changed against the base build
| Part | Base build | This model |
|---|---|---|
| Decoder linear layers | int4 g128 symmetric (AutoRound) | unchanged (65 of 67 files byte-identical) |
lm_head |
int8 RTN, rel. error 0.0064 | int4 GPTQ g128 symmetric, rel. error 0.1407 |
| MTP module (8 linear layers) | int8 RTN, rel. error 0.0066–0.0153 | int4 GPTQ g128 symmetric, rel. error 0.141–0.170 |
| MTP draft head | 40,960 rows of the int8 lm_head |
40,960 rows of the int4 lm_head, same id list |
embed_tokens, vision tower, norms |
int8 / BF16 | unchanged |
| Size on disk | 16.7 GB | 15.8 GB |
Relative error is ||dequant - w|| / ||w|| against the BF16 source weights. GPTQ uses
300,000 hidden states captured from this model's own generations, so it minimizes the
output error on real activations rather than the weight error that this number shows.
On those activations the KL divergence to the BF16 lm_head is 0.00251 for GPTQ
int4, against 0.00558 for round-to-nearest int4. All groups are symmetric (no zero
points).
The MTP module was evaluated on a held-out split of the captured data (greedy chain simulation, 2 draft positions): 2.269 tokens per step, first-position top-1 agreement 0.777, acceptance 0.845, n = 107,063 positions.
The draft vocabulary is the top 40,960 ids of 6,761 generations by this model (4.59 M
output tokens, 3,119 of them with thinking on; prompts: English chat, code, Danish
instructions and reasoning, GSM8K train). On the held-out 10 % of those generations it
covers 97.68 % of tokens; the generic list that the official
Qwen3.8-27B-W4A16-AutoRound checkpoint uses covers 98.18 %. The draft head only
matters for SPEC=mtp.
How it was made
drafter/collect_prompts.py,drafter/gen_data.py,drafter/capture.py: generate with this model and capture its final hidden states.drafter/gptq_lm_head.py --bits 4 --calib-rows 300000: GPTQ int4lm_head.drafter/train_mtp.py --eval-only 1 --dump-hessians, thendrafter/requant_mtp_gptq.py --bits 4: GPTQ int4 MTP module.prepare/build_draft_vocab.py: the draft head from the own-output id list.
The tools are in HyperQwen (drafter/, prepare/).
Properties of the BF16 source
These come from the source and were not measured again on this quantization:
- Refusals (d0xin, fixed 100-prompt set): 88 direct answers, 10 answers with a safety note, 2 other failures, 0 refusals.
- Capability (d0xin, fixed 298-example set against the original Swift BF16): 39.93 % against 38.26 % combined, +1.68 points, McNemar p = 0.44, no measurable loss.
- Method: rank-1 directional residual-stream ablation at layer 38; 131 of 1,199 tensors changed, vision tensors unchanged. Details in the source repository.
- Shorter reasoning (UkisAI): our GSM8K run shows 355 answer tokens on average against 379 for the official AutoRound base, with thinking off.
Measured results (single RTX 3090, 24 GB)
All numbers use the HyperQwen stack on vLLM 0.28.0, WSL2 + Docker, thinking off (protocol v2).
Single-user decode (bench/run_benchmarks.sh single), 2026-09-23, card at 250 W,
SPEC=dflash2 (DFlash2 drafter, 7 draft tokens), PREFIX_CACHE=1, KV_MEM=4529848320
(4.22 GiB), MAX_LEN=49152. End-to-end throughput in tok/s, sampling T=default / T=0.
One run per checkpoint after a warm-up request.
| Concurrency | This model | Base build | Official AutoRound base |
|---|---|---|---|
| C1 | 123.5 / 121.7 | 118.0 / 116.7 | 116.6 / 118.9 |
| C2 | 175.5 / 190.2 | 169.5 / 178.5 | 175.7 / 169.6 |
| C4 | 242.9 / 248.8 | 235.8 / 237.1 | 229.4 / 237.6 |
| C8 | 214.4 / 254.9 | 227.5 / 245.7 | 223.6 / 207.4 |
C1 accepts 3.67 / 3.75 tokens per verify step, mean time to first token 184 ms. Expect 3–5 % variation between sessions.
Batch serving (bench/run_benchmarks.sh batch --prefill --long), 2026-09-21, card at
350 W, KV=fp8, no speculation, second of two runs.
| Row | This model | Base build |
|---|---|---|
| 64 concurrent, 128 in / 512 out | 1,045.2 tok/s | 948.9 tok/s (+10 %) |
| 64 concurrent, 256 in / 256 out | 751.7 tok/s | 701.4 tok/s (+7 %) |
| Prefill 1,024 tokens | 1,824 tok/s | 1,831 tok/s |
| Prefill 102,400 tokens | 1,065 tok/s | 1,065 tok/s |
| 1 x 100k prompt: TTFT / TPOT | 92.8 s / 24.7 ms | 92.8 s / 25.3 ms |
| 4 x 60k prompts, 1,024 out | 17.4 tok/s | 17.2 tok/s |
Quality (bench/quality_battery.py against the served model): perplexity over
~300-token windows of wikitext-2 (en), fineweb-2 Danish (da) and Python source (code);
GSM8K exact match, first 200 test questions, greedy, thinking off.
| Checkpoint | PPL all (en / da / code) | GSM8K | Mean answer tokens |
|---|---|---|---|
| This model | 8.302 (10.85 / 10.98 / 3.35) | 97.0 % (same in an earlier run) | 355 |
| Base build | 8.261 (10.81 / 10.92 / 3.34) | 98.0 % (same) | 356 |
| Official AutoRound base | 8.186 (10.68 / 10.85 / 3.30) | 94.5 % | 379 |
The int4 heads cost 0.5 % perplexity. With n=200 the GSM8K standard error is about 1.2 points at this accuracy, so the one-point GSM8K difference is within noise.
How to serve
HyperQwen (Linux or WSL2, one 24 GB GPU), in .env:
MODEL=/app/models/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
SPEC=dflash2 # or SPEC=mtp to draft with the int4 MTP head
PREFIX_CACHE=1
then docker compose --profile single up -d. bash verify.sh --no-server checks the
directory before you serve it.
Plain vLLM (0.28 or later) loads the body and heads through compressed-tensors:
vllm serve TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast \
--max-model-len 32768 --reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder
The truncated MTP draft head (mtp.draft_lm_head.* with mtp_draft_vocab_ids.pt) is
read by HyperQwen's patches/qwen3_5-mtp-draft-vocab.patch. It is not tested on
unpatched vLLM.
Sampling defaults (generation_config.json, from Swift): temperature 1.0, top_p 0.95,
top_k 20, min_p 0, repetition_penalty 1.0. For thinking off, Qwen recommends
temperature 0.7, top_p 0.8, presence_penalty 1.5. The chat template accepts
reasoning_effort low, medium and xhigh (default), and maps minimal to low
and high / max to xhigh.
Files
| File | Content |
|---|---|
model-000NN-of-00066.safetensors, model_extra_tensors.safetensors |
weights (extras: MTP module, draft head; lm_head is in file 66) |
model.safetensors.index.json |
tensor-to-file map |
config.json, quantization_config.json |
architecture and quantization groups (lm_head int4, embed_tokens int8, mtp.* int4) |
mtp_draft_vocab_ids.pt, draft_vocab_ids.json |
the draft head's 40,960 token ids (same list, two formats) |
chat_template.jinja |
Qwen3.8 template with reasoning-effort aliases and string tool arguments |
tokenizer.json |
the Qwen3.8 tokenizer, byte-identical to the source |
tokenizer_config.json, generation_config.json, preprocessor_config.json, processor_config.json |
tokenizer, sampling and vision configuration |
LICENSE |
Swift Open License v1.0 (verbatim from UkisAI) |
LICENSE-APACHE-2.0 |
Apache License 2.0 of the base model (verbatim from Qwen) |
NOTICE |
attribution chain, UkisAI's notices and the list of changes |
Safety
This is a refusal-reduced model. It can produce content that the upstream model refuses or handles with more care, and its output can be wrong, offensive or unsafe. Use it for research, development and other lawful purposes. Review its output where an error can cause harm. You are responsible for how you use it and for compliance with the licenses and the law. Fine-tuning should start from the BF16 source, not from these 4-bit weights.
License
The weights are distributed under the Swift Open License v1.0 (LICENSE). It
permits use, modification and redistribution. Commercial use is licensed only while
you, together with every entity that controls, is controlled by or is under common
control with you, had gross revenue below US$1,000,000 in the most recently completed
fiscal year (Section 5). Above that threshold, commercial use needs a Swift Enterprise
License from UkisAI. The base model's Apache License 2.0
is included as LICENSE-APACHE-2.0 (Section 4(e)); NOTICE lists the attribution
chain, UkisAI's notices and every change. This summary is not legal advice; LICENSE
is the binding text.
Original Swift model: UkisAI. Abliteration: d0xin. Quantization and serving preparation: TyroneNel.
- Downloads last month
- 97
Model tree for TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
Base model
Qwen/Qwen3.8-27B
docker model run hf.co/TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast