TyroneNel's picture
Upload Swift-Qwen3.8-27B-Uncensored-W4A16-fast
f1944d3 verified
Raw History Blame Contribute Delete
5.1 kB
Swift-Qwen3.8-27B-Uncensored-W4A16 and Swift-Qwen3.8-27B-Uncensored-W4A16-fast
Quantization and serving preparation: Copyright 2026 TyroneNel
This work incorporates the Swift Contribution, licensed under the Swift Open
License v1.0 (see LICENSE):
Swift-Qwen3.8-27B, Copyright 2026 UkisAI.
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.
Provenance chain:
1. Base Model: Qwen/Qwen3.8-27B
https://huggingface.co/Qwen/Qwen3.8-27B
Copyright 2026 Alibaba Cloud. Licensed under the Apache License,
Version 2.0. See LICENSE-APACHE-2.0.
2. Swift Contribution: ukisai/Swift-Qwen3.8-27b
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
Copyright 2026 UkisAI. Licensed under the Swift Open License v1.0.
See LICENSE. A reasoning-efficiency LoRA adapter merged into the Base
Model weights (see UkisAI's NOTICE and model card).
3. Uncensoring: d0xin/Swift-Qwen3.8-27B-Uncensored-BF16
https://huggingface.co/d0xin/Swift-Qwen3.8-27B-Uncensored-BF16
A Derivative Work of (2): rank-1 directional residual-stream ablation
at layer 38 (131 of 1,199 tensors modified, vision tensors unchanged).
Methodology and validation: ABLITERATION.json,
STRUCTURAL_VALIDATION.json and INTELLIGENCE_VALIDATION.json in that
repository.
4. This work: a Derivative Work of (3). Changes (Apache License 2.0,
Section 4(b), and Swift Open License v1.0 change notice):
- model-*.safetensors, model_extra_tensors.safetensors,
model.safetensors.index.json: the decoder linear layers of (3) are
quantized to 4-bit integers (W4A16, group size 128, symmetric) with
Intel AutoRound (128 calibration samples x 2048 tokens). The vision
tower and the GatedDeltaNet in_proj_a/in_proj_b layers stay BF16.
embed_tokens is quantized to int8 (group 128, symmetric).
Swift-Qwen3.8-27B-Uncensored-W4A16: lm_head and the MTP module's
linear layers are quantized to int8 (group 128, symmetric).
Swift-Qwen3.8-27B-Uncensored-W4A16-fast: lm_head and the MTP
module's linear layers are quantized to int4 (group 128,
symmetric) with GPTQ, calibrated on hidden states captured from
this model's own generations.
Both: a 40,960-row draft head (mtp.draft_lm_head.*) is sliced
from the quantized lm_head for MTP speculative decoding.
- config.json, quantization_config.json: quantization metadata
added. generation_config.json: identical to (3).
- mtp_draft_vocab_ids.pt, draft_vocab_ids.json: added (the token ids
of the draft head, counted over this model's own outputs).
- chat_template.jinja: two changes to the Qwen3.8 template.
(a) Reasoning-effort translation: the OpenAI names "minimal",
"high" and "max" map to the template's low / xhigh levels, and an
unknown value no longer raises an error.
(b) Tool calls whose arguments are a JSON string instead of an
object render as one <parameter=arguments> block instead of
failing.
- README.md: replaced.
- tokenizer.json, tokenizer_config.json, preprocessor_config.json,
processor_config.json: re-saved from (3) by transformers 5 (the
chat template moved to chat_template.jinja); same vocabulary,
merges and special tokens.
Tools: the syv-ai/HyperQwen pipeline (run_quant.sh, prepare/,
drafter/), https://github.com/syv-ai/HyperQwen.
Per Section 4(e) of the Swift Open License v1.0, this distribution includes
a copy of the Base Model License (LICENSE-APACHE-2.0) next to the Swift Open
License v1.0 (LICENSE), because this work incorporates portions of the Base
Model.
Attribution notices from the NOTICE file of ukisai/Swift-Qwen3.8-27b,
reproduced as Section 4(d) of the Swift Open License v1.0 requires:
--------------------------------------------------------------------------
Swift-Qwen3.8-27B
Copyright 2026 UkisAI
UkisAI's contribution (the "Swift Contribution") is licensed under the
Swift Open License v1.0. See LICENSE.
This model is a Derivative Work of Qwen3.8-27B
https://huggingface.co/Qwen/Qwen3.8-27B
Copyright 2026 Alibaba Cloud
Licensed under the Apache License, Version 2.0. See LICENSE-APACHE-2.0.
Changes made by UkisAI (Apache License 2.0, Section 4(b) change notice):
- model-*.safetensors, model.safetensors.index.json: model weights were
fine-tuned by UkisAI (LoRA adapter trained by UkisAI and merged into the
Base Model weights).
- generation_config.json: added "min_p": 0 and "repetition_penalty": 1.0.
- README.md: replaced. ukisai-banner.png and swift-speed-demo.mp4 added.
- All other files (config.json, chat_template.jinja, tokenizer.json,
tokenizer_config.json, vocab.json, merges.txt, preprocessor_config.json,
video_preprocessor_config.json) are unmodified from Qwen3.8-27B and
remain under the Apache License, Version 2.0.
--------------------------------------------------------------------------