TokenAIzer commited on
Commit
9b76ac5
·
verified ·
1 Parent(s): bbbb2df

Add model card

Browse files
Files changed (1) hide show
  1. README.md +187 -0
README.md ADDED
@@ -0,0 +1,187 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: ukisai/Swift-Qwen3.8-27b
3
+ model_name: Swift-Qwen3.8-27b-oQ4-mtp
4
+ library_name: mlx
5
+ pipeline_tag: image-text-to-text
6
+ license: other
7
+ license_name: swift-open-license-1.0
8
+ license_link: "https://huggingface.co/ukisai/Swift-Qwen3.8-27b#license-and-access"
9
+ tags:
10
+ - mlx
11
+ - omlx
12
+ - quantization
13
+ - mixed-precision
14
+ - apple-silicon
15
+ - mtp
16
+ - speculative-decoding
17
+ - qwen
18
+ - vision
19
+ - base_model:quantized:ukisai/Swift-Qwen3.8-27b
20
+ - base_model_size:10B to 100B
21
+ ---
22
+
23
+ # Swift-Qwen3.8-27b-oQ4-mtp — mixed 4/5-bit MLX quant of Swift-Qwen3.8-27b (MTP head kept)
24
+
25
+ Unofficial Apple Silicon quantization of **[ukisai/Swift-Qwen3.8-27b](https://huggingface.co/ukisai/Swift-Qwen3.8-27b)**, produced with **oMLX 0.6.4** (MLX `affine` quantizer, group size 64) on an Apple M5 Max / 128 GB.
26
+
27
+ Two things are preserved on purpose:
28
+
29
+ - **Swift's reasoning efficiency.** Smallest, fastest to load. Best default for coding/agentic work where you want the shortest reasoning and maximum headroom for KV cache.
30
+ - **The MTP head.** The `-mtp` suffix means the multi-token-prediction head from the base checkpoint ships intact (29 tensors, `mtp_num_hidden_layers: 1`), so oMLX can run self-speculative decoding instead of wasting the weights.
31
+
32
+ I am not affiliated with UkisAI. All upstream weights, benchmarks and license terms belong to UkisAI, and **the upstream license governs this repository too** (see [License](#license)).
33
+
34
+ > Format note: these are **MLX safetensors**, not GGUF. They will not load in llama.cpp / Ollama / LM Studio. Use oMLX or MLX runtimes.
35
+
36
+ ## Pick a variant
37
+
38
+ | | oQ4-mtp | oQ6-mtp |
39
+ |---|---|---|
40
+ | Weights on disk | 15.81 GiB (16.97 GB), 4 shards | 22.09 GiB (23.72 GB), 5 shards |
41
+ | Weight precision | mixed 4/5-bit, group 64 | mixed 6/8-bit, group 64 |
42
+ | Effective bits / quantized weight | 4.68 | 6.66 |
43
+ | Effective bits / parameter (whole repo) | 4.89 | 6.83 |
44
+ | Minimum Apple Silicon RAM | 24 GB (context ≲ 32k) | 32 GB (short context) |
45
+ | Comfortable | 32 GB+ | 48 GB+ |
46
+ | Best for | memory-bound, long agentic sessions | max fidelity, math, vision |
47
+
48
+ Smallest, fastest to load. Best default for coding/agentic work where you want the shortest reasoning and maximum headroom for KV cache.
49
+
50
+ ## What is inside (read straight from the shipped `config.json`)
51
+
52
+ | Field | Value |
53
+ |---|---|
54
+ | Architecture | `Qwen3_5ForConditionalGeneration` (`model_type: qwen3_5`) |
55
+ | Parameters | 27.78 B total, 27.27 B quantized (98.1%) |
56
+ | Text layers / hidden | 64 layers, `hidden_size` 5120, `intermediate_size` 17408 |
57
+ | Attention | hybrid: 1 full-attention layer every 4 (`full_attention_interval: 4`), 24 heads / 4 KV, `head_dim` 256, `attn_output_gate: true`; the rest are gated linear-attention (`linear_attn`) |
58
+ | Context | `max_position_embeddings: 262144` |
59
+ | Vocab | 248,320 (tokenizer and `chat_template.jinja` copied from upstream, unchanged) |
60
+ | MTP | `mtp_num_hidden_layers: 1`, `mtp_use_dedicated_embeddings: false`, weights included |
61
+ | Vision | Qwen vision tower kept in **BF16** (~0.92 GB), depth 27, patch 16, spatial merge 2 |
62
+ | Metadata | `{"format": "mlx"}` in every safetensors header |
63
+
64
+ ## Quantization recipe
65
+
66
+ Precision is mixed per module and recorded verbatim in `config.json` → `quantization_config`, so any MLX loader reproduces the layout without guessing:
67
+
68
+ - default: **mixed**, `group_size: 64`, `mode: affine`
69
+ - 339 modules @ 4-bit + 166 modules bumped to 5-bit
70
+ - bumped modules: early layers (`linear_attn.in_proj_a/b/z`, `linear_attn.out_proj`, `mlp.down_proj`, some `self_attn.k_proj/o_proj`)
71
+ - **never quantized:** vision tower (BF16), all `scales`/`biases` (1.68 GB BF16), norms, `A_log`, `dt_bias`, convolutions (≈5 MB), MTP non-linear weights (128 MB)
72
+
73
+ ## Requirements
74
+
75
+ - Apple Silicon, macOS 15+ (built and smoke-tested on M5 Max, 128 GB)
76
+ - oMLX **≥ 0.6.4**, or a recent `mlx` / `mlx-lm` / `mlx-vlm` build with `qwen3_5` support
77
+
78
+ ## Usage
79
+
80
+ ### oMLX (the runtime these were made for)
81
+
82
+ ```bash
83
+ # 1. drop the folder into the oMLX model dir
84
+ git clone https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp ~/.omlx/models/Swift-Qwen3.8-27b-oQ4-mtp
85
+
86
+ # 2. start the multi-model server (model id = folder name)
87
+ omlx serve --model-dir ~/.omlx/models --port 8000
88
+
89
+ # 3. talk to it
90
+ curl -s http://127.0.0.1:8000/v1/chat/completions \
91
+ -H 'Content-Type: application/json' \
92
+ -d '{"model": "Swift-Qwen3.8-27b-oQ4-mtp",
93
+ "messages": [{"role": "user", "content": "Explain speculative decoding in two sentences."}]}'
94
+ ```
95
+
96
+ In oMLX model settings, enable the speculative head and the matching reasoning parser:
97
+
98
+ ```json
99
+ {
100
+ "mtp_enabled": true,
101
+ "reasoning_parser": "qwen_3_5",
102
+ "max_context_window": 262144,
103
+ "model_type_override": "vlm"
104
+ }
105
+ ```
106
+
107
+ ### MLX directly
108
+
109
+ ```bash
110
+ pip install -U mlx-lm mlx-vlm
111
+ python -m mlx_lm.server --model suzu89/Swift-Qwen3.8-27b-oQ4-mtp --port 8000
112
+ ```
113
+
114
+ Only oMLX 0.6.4 is verified by me; if you get `mlx_lm` running this architecture, please open an issue and I will document it.
115
+
116
+ ### Not supported
117
+
118
+ `llama.cpp`, GGUF, vLLM and SGLang paths in the upstream card do not apply here — this repo has no GGUF and no PyTorch weights. For BF16/server deployments use `ukisai/Swift-Qwen3.8-27b`.
119
+
120
+ ## Recommended sampling
121
+
122
+ Shipped `generation_config.json` (unchanged from upstream) is the tuning target for thinking mode:
123
+
124
+ ```
125
+ temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · repetition_penalty 1.0
126
+ eos_token_id [248046, 248044]
127
+ ```
128
+
129
+ The upstream chat template supports tool calling, image/video inputs and an `enable_thinking` switch, so you can trade reasoning length per request; upstream reports Swift's token savings hold at `xhigh`, `medium` and `low` reasoning effort.
130
+
131
+ ## Benchmarks
132
+
133
+ I publish no numbers I have not measured myself. This table is the honest state of the repository:
134
+
135
+ | Benchmark | oQ4-mtp | oQ6-mtp | BF16 upstream (reference) |
136
+ |---|---|---|---|
137
+ | GPQA-Diamond | not measured | not measured | 88.28% |
138
+ | AIME 2026 | not measured | not measured | 94.00% |
139
+ | LiveCodeBench v6 | not measured | not measured | 81.55% |
140
+ | IFBench | not measured | not measured | 71.80% |
141
+
142
+ Upstream Swift vs Qwen3.8-27B results (GPQA-Diamond −0.1 pt for ~41% fewer mean thinking tokens, ~1.95× faster) are reported by UkisAI in the [base model card](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) and are **not** measurements of these quantized weights. What you should realistically expect from a quantization of a "think less" fine-tune: token savings largely survive (they come from behaviour, not precision), while the hardest math splits and long-horizon tool chains degrade slightly — most visibly at 4-bit.
143
+
144
+ Measured throughput/acceptance-length data and issue reports (especially "quant X broke task Y") are welcome and will be merged into this table.
145
+
146
+ ## Known caveats
147
+
148
+ - Quantization is lossy: expect small regressions versus BF16, largest on competition math and very long agentic traces. Try `oQ6-mtp` before filing a bug.
149
+ - `oQ4-mtp` can amplify repetition on degenerate loops; keep `repetition_penalty` at 1.0 first and only then nudge it.
150
+ - Vision works through the BF16 tower, but I have not benchmarked VQA accuracy post-quantization.
151
+ - 262k context is the architecture's limit, not a promise: keep KV cache within your memory budget or the system swaps.
152
+ - MTP decoding only helps when the speculative draft is enabled in the runtime; without it you pay for the head and get nothing.
153
+
154
+ ## License
155
+
156
+ **This repository is distributed under the [Swift Open License v1.0](https://huggingface.co/ukisai/Swift-Qwen3.8-27b#license-and-access).** A quantization is a derivative work: it inherits the upstream terms in full and cannot be released under a more permissive license.
157
+
158
+ - Free personal, research, educational, evaluation and commercial use for individuals and organizations with annual recurring revenue (including affiliates) **up to US$1,000,000**.
159
+ - Above that threshold, commercial use requires a separate **Swift Enterprise License** from UkisAI.
160
+ - The base Qwen3.8 checkpoint and the ThinkingCap-Qwen3.6-27B transfer component (BottleCap AI) contribute their own terms, which apply to you as well — read the `LICENSE` files in [`ukisai/Swift-Qwen3.8-27b`](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) and the BottleCap repository before commercial deployment.
161
+ - Keep this attribution, the upstream citation and the `base_model` metadata intact when you redistribute.
162
+
163
+ ## Citation
164
+
165
+ ```bibtex
166
+ @misc{swift-qwen3.8-27b-mlx-quants,
167
+ title = {Swift-Qwen3.8-27b-oQ4-mtp}: oMLX/MLX quantization of Swift-Qwen3.8-27B with MTP head retained,
168
+ author = {suzu89},
169
+ year = {2026},
170
+ howpublished = {\url{https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp}},
171
+ note = {Unofficial quantization of ukisai/Swift-Qwen3.8-27b}
172
+ }
173
+
174
+ @misc{swift-qwen3.8-27b,
175
+ title = {Swift-Qwen3.8-27B},
176
+ author = {UkisAI},
177
+ year = {2026},
178
+ url = {https://huggingface.co/ukisai/Swift-Qwen3.8-27b}
179
+ }
180
+ ```
181
+
182
+ ## Acknowledgements
183
+
184
+ - **UkisAI** for Swift-Qwen3.8-27B and the public evaluation harness.
185
+ - **Qwen team** for the Qwen3.8-27B base model.
186
+ - **BottleCap AI** for the ThinkingCap-Qwen3.6-27B transfer component used upstream.
187
+ - **oMLX** for the Apple Silicon server and quantizer that made these builds possible.