liamwh commited on
Commit
5e9ab3f
·
verified ·
1 Parent(s): f6d2344

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +82 -61
README.md CHANGED
@@ -1,75 +1,96 @@
1
  ---
2
  license: other
3
  license_name: swift-open-license-1.0
4
- library_name: transformers
5
- pipeline_tag: image-text-to-text
6
- base_model: ukisai/Swift-Qwen3.8-27b
 
 
7
  tags:
8
  - qwen
9
- - qwen3.5
10
  - qwen3.8
11
- - finetune
12
- - awq
13
- - int4
14
  - w4a16
 
15
  - compressed-tensors
16
- - lmdeploy
 
 
 
 
17
  ---
18
 
19
- # Swift-Qwen3.8-27B-W4A16-AWQ
20
-
21
- W4A16 (4-bit weights, 16-bit activations) AWQ compressed-tensors quantization
22
- of [`ukisai/Swift-Qwen3.8-27b`](https://huggingface.co/ukisai/Swift-Qwen3.8-27b).
23
-
24
- ## Quantization method
25
-
26
- - **Scheme:** `W4A16_ASYM` — 4-bit asymmetric per-group quantization (group size 128) of all `Linear` weights, stored in the compressed-tensors pack-quantized format (`weight_packed` / `weight_scale` / `weight_zero_point` / `weight_shape`), which LMDeploy `turbomind` auto-detects and loads natively (including the MTP head and vision tower, which stay BF16).
27
- - **Tooling:** [llmcompressor](https://github.com/vllm-project/llmcompressor) one-shot offline quantization with CPU offloading (`compressed_tensors.offload.load_offloaded_model`), so the full-precision source fits on a 2×16 GB VRAM setup.
28
- - **AWQ activation smoothing:** `AWQModifier` with the layer-scoped hybrid-attention mappings from `build_hybrid_attention_mappings` — full-attention `input_layernorm` → `self_attn.q/k/v`, `post_attention_layernorm` → `mlp.gate/up`, and `mlp.up_proj` → `mlp.down_proj`, with `duo_scaling="both"` and CPU offload, followed by W4A16 quantization. This layer-scoped recipe is required for hybrid-attention (Qwen3.5-family) architectures — grouped-regex smoothing or mismatched mappings corrupt decoding.
29
- - Unquantized (kept BF16): embeddings, `lm_head`, norms, `linear_attn.in_proj_a/b`, the vision tower, and the MTP head.
30
- - **Run command:**
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
31
 
32
  ```bash
33
- CUDA_VISIBLE_DEVICES=1,2 python3 quantize-awq-hybrid.py \
34
- --model_path ./Swift-Qwen3.8-27b \
35
- --quant_path ./Swift-Qwen3.8-27b-W4A16-AWQ \
36
- --offload_dir ./Swift-Qwen3.8-27b-W4A16-AWQ
37
  ```
38
 
39
- (Default calibration: UltraChat 200k `train_sft`.)
40
-
41
- ## Recommended parameters
42
-
43
- From the [base model](https://huggingface.co/ukisai/Swift-Qwen3.8-27b): temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0, repetition_penalty 1.0.
44
-
45
- ## Usage
46
-
47
- Tested with LMDeploy `turbomind`:
48
-
49
- ```bash
50
- from lmdeploy import pipeline, TurbomindEngineConfig
51
-
52
- pipe = pipeline(
53
- "TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ",
54
- backend_config=TurbomindEngineConfig(
55
- tp=2,
56
- model_format="compressed-tensors",
57
- language_model_only=True,
58
- ),
59
- )
60
- print(pipe("Hello, who are you?").text)
61
- ```
62
-
63
- ## License and access
64
-
65
- Swift weights are distributed through gated access under the **Swift Open License v1.0**;
66
- this quantization inherits that license from the base model. Personal, research,
67
- educational, evaluation, and commercial use are free for individuals and organizations
68
- with annual recurring revenue, including affiliates, of up to US$1,000,000. Above that
69
- threshold, commercial use requires a separate **Swift Enterprise License**. Contact
70
- [UkisAI](https://ukisai.com/contact) for terms.
71
-
72
- ## Files
73
-
74
- - `quantize-awq-hybrid.py` — the script used to produce this quantization (CPU-offloaded, DDP/torchrun, produces properly numbered `-of-N` shards). `--offload_dir` selects where per-rank CPU offload temp folders live (defaults to the current working directory).
75
- - `model-nonquant.safetensors` — unquantized tensors (`mtp.*` and `model.visual.*`) preserved BF16 so the full model architecture is loadable.
 
1
  ---
2
  license: other
3
  license_name: swift-open-license-1.0
4
+ license_link: LICENSE
5
+ library_name: vllm
6
+ base_model:
7
+ - ukisai/Swift-Qwen3.8-27b
8
+ - TheUnderscore/Swift-Qwen3.8-27b-W4A16-AWQ
9
  tags:
10
  - qwen
 
11
  - qwen3.8
12
+ - swift
 
 
13
  - w4a16
14
+ - gptq
15
  - compressed-tensors
16
+ - speculative-decoding
17
+ - mtp
18
+ - rtx-3090
19
+ - syv
20
+ pipeline_tag: text-generation
21
  ---
22
 
23
+ # Swift-Qwen3.8-27B-W4A16-syv-fast
24
+
25
+ The [syv-ai/qwen38-27b-rtx3090](https://github.com/syv-ai/qwen38-27b-rtx3090)
26
+ **fast-variant serving shape** of
27
+ [ukisai/Swift-Qwen3.8-27b](https://huggingface.co/ukisai/Swift-Qwen3.8-27b)
28
+ (the "reduced reasoning" finetune of Qwen3.8-27B), built entirely from
29
+ Swift's own weights and outputs:
30
+
31
+ - **Body**: TheUnderscore's W4A16 asymmetric-AWQ g128 compressed-tensors
32
+ quant of Swift, carried through unmodified.
33
+ - **lm_head + MTP**: int4 GPTQ (g128, symmetric) calibrated on **Swift's own
34
+ hidden states** — 300k rows captured through the syv `drafter/` pipeline.
35
+ lm_head KL to the bf16 head: **0.00234** (RTN: 0.00707; the official syv
36
+ Qwen fast variant shipped at 0.0029). MTP relative errors 0.146–0.176,
37
+ inside the range of the shipped syv fast variant.
38
+ - **Embeddings**: int8 g128, requantized from Swift's own weights.
39
+ - **Draft vocab**: counted over 4.23M tokens of **Swift's own outputs** on a
40
+ coding-agent-weighted prompt corpus (35% Rust, 16% TypeScript, 11%
41
+ debugging, 10% code-edit, 8% agent/tool-use, 10% architecture, 7%
42
+ technical reasoning, 3% general; 61% thinking-on). Swift emits far fewer
43
+ distinct tokens than base Qwen (~25.9k vs ~54k), so the draft head is
44
+ 25,879 rows, not the base model's 40,960.
45
+
46
+ Held-out coverage (10% of sequences, never counted against):
47
+
48
+ | draft vocab | all tokens | code sources |
49
+ |---|---|---|
50
+ | **Swift-derived (this model)** | **99.81%** | **99.86%** |
51
+ | base-Qwen 40k list | 96.69% | 96.67% |
52
+
53
+ ## Serving
54
+
55
+ Designed for the syv single-user stack (vLLM 0.28.0 + the repo's patch
56
+ series — the MTP draft-vocab patch is required for the 25,879-row head;
57
+ stock vLLM will not use it):
58
 
59
  ```bash
60
+ MODEL=/path/to/Swift-Qwen3.8-27B-W4A16-syv-fast \
61
+ SPEC=mtp PREFIX_CACHE=1 CTX=long MAX_LEN=114688 \
62
+ bash single-user/start_qwen.sh # from the syv checkout
 
63
  ```
64
 
65
+ The syv `verify.sh` will report one false FAIL on this dir (it asserts an
66
+ int8 lm_head; this is int4 by construction, like the official fast variant).
67
+
68
+ ## Measured (RTX 3090, syv stack, MTP + prefix caching, 114,688 context)
69
+
70
+ | | decode tok/s | MTP acceptance | tok/step |
71
+ |---|---|---|---|
72
+ | **this model** | 98.4 | **0.660** | 2.98 |
73
+ | Swift + int8 heads, base draft vocab | 94.0 | 0.630 | 2.89 |
74
+ | official syv Qwen fast variant | 98.2 | 0.634 | 2.90 |
75
+
76
+ Quality battery (identical prompts): 8/9, matching the int8 build — tool
77
+ calling, strict JSON, streaming and the qwen3 reasoning parser all clean.
78
+ Swift's reasoning-termination behaviour is preserved (GPTQ heads verified
79
+ not to shift it, including under greedy decoding).
80
+
81
+ ## Provenance & licence
82
+
83
+ Attribution chain: **Alibaba Cloud** Qwen3.8-27B (Apache-2.0, included as
84
+ `LICENSE-APACHE-2.0`) → **UkisAI** Swift finetune (Swift Open License v1.0,
85
+ included as `LICENSE`) → **TheUnderscore** W4A16-AWQ body (same licence) →
86
+ this repository's int4-GPTQ heads, int8 embeddings and Swift-derived draft
87
+ vocabulary (quantisation and calibration by **liamwh**, using the
88
+ [syv-ai/qwen38-27b-rtx3090](https://github.com/syv-ai/qwen38-27b-rtx3090)
89
+ `drafter/` pipeline with a coding-agent-weighted Swift corpus). Distributed
90
+ under the Swift Open License v1.0; commercial use above its revenue
91
+ threshold requires the Swift Enterprise License.
92
+
93
+ Rebuild from scratch with the syv checkout's `drafter/` pipeline
94
+ (`gen_data.py` → `capture.py` → GPTQ heads → `build_draft_vocab.py`) over a
95
+ Swift-weighted prompt corpus; `swift_draft_vocab_ids.json` here is the
96
+ exact id list this model serves.