Image-Text-to-Text
Transformers
Safetensors
English
Chinese
qwen3_5
qwen3.8
swift
abliterated
uncensored
w4a16
4-bit precision
int4
int8
gptq
compressed-tensors
autoround
vllm
marlin
mtp
speculative-decoding
efficient-thinking
reasoning
function-calling
vision-language
rtx-3090
conversational
Instructions to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast") model = AutoModelForMultimodalLM.from_pretrained("TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
- SGLang
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast with Docker Model Runner:
docker model run hf.co/TyroneNel/Swift-Qwen3.8-27B-Uncensored-W4A16-fast
File size: 5,100 Bytes
f1944d3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 | Swift-Qwen3.8-27B-Uncensored-W4A16 and Swift-Qwen3.8-27B-Uncensored-W4A16-fast
Quantization and serving preparation: Copyright 2026 TyroneNel
This work incorporates the Swift Contribution, licensed under the Swift Open
License v1.0 (see LICENSE):
Swift-Qwen3.8-27B, Copyright 2026 UkisAI.
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.
Provenance chain:
1. Base Model: Qwen/Qwen3.8-27B
https://huggingface.co/Qwen/Qwen3.8-27B
Copyright 2026 Alibaba Cloud. Licensed under the Apache License,
Version 2.0. See LICENSE-APACHE-2.0.
2. Swift Contribution: ukisai/Swift-Qwen3.8-27b
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
Copyright 2026 UkisAI. Licensed under the Swift Open License v1.0.
See LICENSE. A reasoning-efficiency LoRA adapter merged into the Base
Model weights (see UkisAI's NOTICE and model card).
3. Uncensoring: d0xin/Swift-Qwen3.8-27B-Uncensored-BF16
https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-BF16
A Derivative Work of (2): rank-1 directional residual-stream ablation
at layer 38 (131 of 1,199 tensors modified, vision tensors unchanged).
Methodology and validation: ABLITERATION.json,
STRUCTURAL_VALIDATION.json and INTELLIGENCE_VALIDATION.json in that
repository.
4. This work: a Derivative Work of (3). Changes (Apache License 2.0,
Section 4(b), and Swift Open License v1.0 change notice):
- model-*.safetensors, model_extra_tensors.safetensors,
model.safetensors.index.json: the decoder linear layers of (3) are
quantized to 4-bit integers (W4A16, group size 128, symmetric) with
Intel AutoRound (128 calibration samples x 2048 tokens). The vision
tower and the GatedDeltaNet in_proj_a/in_proj_b layers stay BF16.
embed_tokens is quantized to int8 (group 128, symmetric).
Swift-Qwen3.8-27B-Uncensored-W4A16: lm_head and the MTP module's
linear layers are quantized to int8 (group 128, symmetric).
Swift-Qwen3.8-27B-Uncensored-W4A16-fast: lm_head and the MTP
module's linear layers are quantized to int4 (group 128,
symmetric) with GPTQ, calibrated on hidden states captured from
this model's own generations.
Both: a 40,960-row draft head (mtp.draft_lm_head.*) is sliced
from the quantized lm_head for MTP speculative decoding.
- config.json, quantization_config.json: quantization metadata
added. generation_config.json: identical to (3).
- mtp_draft_vocab_ids.pt, draft_vocab_ids.json: added (the token ids
of the draft head, counted over this model's own outputs).
- chat_template.jinja: two changes to the Qwen3.8 template.
(a) Reasoning-effort translation: the OpenAI names "minimal",
"high" and "max" map to the template's low / xhigh levels, and an
unknown value no longer raises an error.
(b) Tool calls whose arguments are a JSON string instead of an
object render as one <parameter=arguments> block instead of
failing.
- README.md: replaced.
- tokenizer.json, tokenizer_config.json, preprocessor_config.json,
processor_config.json: re-saved from (3) by transformers 5 (the
chat template moved to chat_template.jinja); same vocabulary,
merges and special tokens.
Tools: the syv-ai/HyperQwen pipeline (run_quant.sh, prepare/,
drafter/), https://github.com/syv-ai/HyperQwen.
Per Section 4(e) of the Swift Open License v1.0, this distribution includes
a copy of the Base Model License (LICENSE-APACHE-2.0) next to the Swift Open
License v1.0 (LICENSE), because this work incorporates portions of the Base
Model.
Attribution notices from the NOTICE file of ukisai/Swift-Qwen3.8-27b,
reproduced as Section 4(d) of the Swift Open License v1.0 requires:
--------------------------------------------------------------------------
Swift-Qwen3.8-27B
Copyright 2026 UkisAI
UkisAI's contribution (the "Swift Contribution") is licensed under the
Swift Open License v1.0. See LICENSE.
This model is a Derivative Work of Qwen3.8-27B
https://huggingface.co/Qwen/Qwen3.8-27B
Copyright 2026 Alibaba Cloud
Licensed under the Apache License, Version 2.0. See LICENSE-APACHE-2.0.
Changes made by UkisAI (Apache License 2.0, Section 4(b) change notice):
- model-*.safetensors, model.safetensors.index.json: model weights were
fine-tuned by UkisAI (LoRA adapter trained by UkisAI and merged into the
Base Model weights).
- generation_config.json: added "min_p": 0 and "repetition_penalty": 1.0.
- README.md: replaced. ukisai-banner.png and swift-speed-demo.mp4 added.
- All other files (config.json, chat_template.jinja, tokenizer.json,
tokenizer_config.json, vocab.json, merges.txt, preprocessor_config.json,
video_preprocessor_config.json) are unmodified from Qwen3.8-27B and
remain under the Apache License, Version 2.0.
--------------------------------------------------------------------------
|