Image-Text-to-Text
MLX
Safetensors
qwen3_5
omlx
quantization
mixed-precision
apple-silicon
mtp
speculative-decoding
qwen
vision
base_model_size:10B to 100B
conversational
4-bit precision
Instructions to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp") config = load_config("TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
TokenAIzer commited on
Add model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,187 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: ukisai/Swift-Qwen3.8-27b
|
| 3 |
+
model_name: Swift-Qwen3.8-27b-oQ4-mtp
|
| 4 |
+
library_name: mlx
|
| 5 |
+
pipeline_tag: image-text-to-text
|
| 6 |
+
license: other
|
| 7 |
+
license_name: swift-open-license-1.0
|
| 8 |
+
license_link: "https://huggingface.co/ukisai/Swift-Qwen3.8-27b#license-and-access"
|
| 9 |
+
tags:
|
| 10 |
+
- mlx
|
| 11 |
+
- omlx
|
| 12 |
+
- quantization
|
| 13 |
+
- mixed-precision
|
| 14 |
+
- apple-silicon
|
| 15 |
+
- mtp
|
| 16 |
+
- speculative-decoding
|
| 17 |
+
- qwen
|
| 18 |
+
- vision
|
| 19 |
+
- base_model:quantized:ukisai/Swift-Qwen3.8-27b
|
| 20 |
+
- base_model_size:10B to 100B
|
| 21 |
+
---
|
| 22 |
+
|
| 23 |
+
# Swift-Qwen3.8-27b-oQ4-mtp — mixed 4/5-bit MLX quant of Swift-Qwen3.8-27b (MTP head kept)
|
| 24 |
+
|
| 25 |
+
Unofficial Apple Silicon quantization of **[ukisai/Swift-Qwen3.8-27b](https://huggingface.co/ukisai/Swift-Qwen3.8-27b)**, produced with **oMLX 0.6.4** (MLX `affine` quantizer, group size 64) on an Apple M5 Max / 128 GB.
|
| 26 |
+
|
| 27 |
+
Two things are preserved on purpose:
|
| 28 |
+
|
| 29 |
+
- **Swift's reasoning efficiency.** Smallest, fastest to load. Best default for coding/agentic work where you want the shortest reasoning and maximum headroom for KV cache.
|
| 30 |
+
- **The MTP head.** The `-mtp` suffix means the multi-token-prediction head from the base checkpoint ships intact (29 tensors, `mtp_num_hidden_layers: 1`), so oMLX can run self-speculative decoding instead of wasting the weights.
|
| 31 |
+
|
| 32 |
+
I am not affiliated with UkisAI. All upstream weights, benchmarks and license terms belong to UkisAI, and **the upstream license governs this repository too** (see [License](#license)).
|
| 33 |
+
|
| 34 |
+
> Format note: these are **MLX safetensors**, not GGUF. They will not load in llama.cpp / Ollama / LM Studio. Use oMLX or MLX runtimes.
|
| 35 |
+
|
| 36 |
+
## Pick a variant
|
| 37 |
+
|
| 38 |
+
| | oQ4-mtp | oQ6-mtp |
|
| 39 |
+
|---|---|---|
|
| 40 |
+
| Weights on disk | 15.81 GiB (16.97 GB), 4 shards | 22.09 GiB (23.72 GB), 5 shards |
|
| 41 |
+
| Weight precision | mixed 4/5-bit, group 64 | mixed 6/8-bit, group 64 |
|
| 42 |
+
| Effective bits / quantized weight | 4.68 | 6.66 |
|
| 43 |
+
| Effective bits / parameter (whole repo) | 4.89 | 6.83 |
|
| 44 |
+
| Minimum Apple Silicon RAM | 24 GB (context ≲ 32k) | 32 GB (short context) |
|
| 45 |
+
| Comfortable | 32 GB+ | 48 GB+ |
|
| 46 |
+
| Best for | memory-bound, long agentic sessions | max fidelity, math, vision |
|
| 47 |
+
|
| 48 |
+
Smallest, fastest to load. Best default for coding/agentic work where you want the shortest reasoning and maximum headroom for KV cache.
|
| 49 |
+
|
| 50 |
+
## What is inside (read straight from the shipped `config.json`)
|
| 51 |
+
|
| 52 |
+
| Field | Value |
|
| 53 |
+
|---|---|
|
| 54 |
+
| Architecture | `Qwen3_5ForConditionalGeneration` (`model_type: qwen3_5`) |
|
| 55 |
+
| Parameters | 27.78 B total, 27.27 B quantized (98.1%) |
|
| 56 |
+
| Text layers / hidden | 64 layers, `hidden_size` 5120, `intermediate_size` 17408 |
|
| 57 |
+
| Attention | hybrid: 1 full-attention layer every 4 (`full_attention_interval: 4`), 24 heads / 4 KV, `head_dim` 256, `attn_output_gate: true`; the rest are gated linear-attention (`linear_attn`) |
|
| 58 |
+
| Context | `max_position_embeddings: 262144` |
|
| 59 |
+
| Vocab | 248,320 (tokenizer and `chat_template.jinja` copied from upstream, unchanged) |
|
| 60 |
+
| MTP | `mtp_num_hidden_layers: 1`, `mtp_use_dedicated_embeddings: false`, weights included |
|
| 61 |
+
| Vision | Qwen vision tower kept in **BF16** (~0.92 GB), depth 27, patch 16, spatial merge 2 |
|
| 62 |
+
| Metadata | `{"format": "mlx"}` in every safetensors header |
|
| 63 |
+
|
| 64 |
+
## Quantization recipe
|
| 65 |
+
|
| 66 |
+
Precision is mixed per module and recorded verbatim in `config.json` → `quantization_config`, so any MLX loader reproduces the layout without guessing:
|
| 67 |
+
|
| 68 |
+
- default: **mixed**, `group_size: 64`, `mode: affine`
|
| 69 |
+
- 339 modules @ 4-bit + 166 modules bumped to 5-bit
|
| 70 |
+
- bumped modules: early layers (`linear_attn.in_proj_a/b/z`, `linear_attn.out_proj`, `mlp.down_proj`, some `self_attn.k_proj/o_proj`)
|
| 71 |
+
- **never quantized:** vision tower (BF16), all `scales`/`biases` (1.68 GB BF16), norms, `A_log`, `dt_bias`, convolutions (≈5 MB), MTP non-linear weights (128 MB)
|
| 72 |
+
|
| 73 |
+
## Requirements
|
| 74 |
+
|
| 75 |
+
- Apple Silicon, macOS 15+ (built and smoke-tested on M5 Max, 128 GB)
|
| 76 |
+
- oMLX **≥ 0.6.4**, or a recent `mlx` / `mlx-lm` / `mlx-vlm` build with `qwen3_5` support
|
| 77 |
+
|
| 78 |
+
## Usage
|
| 79 |
+
|
| 80 |
+
### oMLX (the runtime these were made for)
|
| 81 |
+
|
| 82 |
+
```bash
|
| 83 |
+
# 1. drop the folder into the oMLX model dir
|
| 84 |
+
git clone https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp ~/.omlx/models/Swift-Qwen3.8-27b-oQ4-mtp
|
| 85 |
+
|
| 86 |
+
# 2. start the multi-model server (model id = folder name)
|
| 87 |
+
omlx serve --model-dir ~/.omlx/models --port 8000
|
| 88 |
+
|
| 89 |
+
# 3. talk to it
|
| 90 |
+
curl -s http://127.0.0.1:8000/v1/chat/completions \
|
| 91 |
+
-H 'Content-Type: application/json' \
|
| 92 |
+
-d '{"model": "Swift-Qwen3.8-27b-oQ4-mtp",
|
| 93 |
+
"messages": [{"role": "user", "content": "Explain speculative decoding in two sentences."}]}'
|
| 94 |
+
```
|
| 95 |
+
|
| 96 |
+
In oMLX model settings, enable the speculative head and the matching reasoning parser:
|
| 97 |
+
|
| 98 |
+
```json
|
| 99 |
+
{
|
| 100 |
+
"mtp_enabled": true,
|
| 101 |
+
"reasoning_parser": "qwen_3_5",
|
| 102 |
+
"max_context_window": 262144,
|
| 103 |
+
"model_type_override": "vlm"
|
| 104 |
+
}
|
| 105 |
+
```
|
| 106 |
+
|
| 107 |
+
### MLX directly
|
| 108 |
+
|
| 109 |
+
```bash
|
| 110 |
+
pip install -U mlx-lm mlx-vlm
|
| 111 |
+
python -m mlx_lm.server --model suzu89/Swift-Qwen3.8-27b-oQ4-mtp --port 8000
|
| 112 |
+
```
|
| 113 |
+
|
| 114 |
+
Only oMLX 0.6.4 is verified by me; if you get `mlx_lm` running this architecture, please open an issue and I will document it.
|
| 115 |
+
|
| 116 |
+
### Not supported
|
| 117 |
+
|
| 118 |
+
`llama.cpp`, GGUF, vLLM and SGLang paths in the upstream card do not apply here — this repo has no GGUF and no PyTorch weights. For BF16/server deployments use `ukisai/Swift-Qwen3.8-27b`.
|
| 119 |
+
|
| 120 |
+
## Recommended sampling
|
| 121 |
+
|
| 122 |
+
Shipped `generation_config.json` (unchanged from upstream) is the tuning target for thinking mode:
|
| 123 |
+
|
| 124 |
+
```
|
| 125 |
+
temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · repetition_penalty 1.0
|
| 126 |
+
eos_token_id [248046, 248044]
|
| 127 |
+
```
|
| 128 |
+
|
| 129 |
+
The upstream chat template supports tool calling, image/video inputs and an `enable_thinking` switch, so you can trade reasoning length per request; upstream reports Swift's token savings hold at `xhigh`, `medium` and `low` reasoning effort.
|
| 130 |
+
|
| 131 |
+
## Benchmarks
|
| 132 |
+
|
| 133 |
+
I publish no numbers I have not measured myself. This table is the honest state of the repository:
|
| 134 |
+
|
| 135 |
+
| Benchmark | oQ4-mtp | oQ6-mtp | BF16 upstream (reference) |
|
| 136 |
+
|---|---|---|---|
|
| 137 |
+
| GPQA-Diamond | not measured | not measured | 88.28% |
|
| 138 |
+
| AIME 2026 | not measured | not measured | 94.00% |
|
| 139 |
+
| LiveCodeBench v6 | not measured | not measured | 81.55% |
|
| 140 |
+
| IFBench | not measured | not measured | 71.80% |
|
| 141 |
+
|
| 142 |
+
Upstream Swift vs Qwen3.8-27B results (GPQA-Diamond −0.1 pt for ~41% fewer mean thinking tokens, ~1.95× faster) are reported by UkisAI in the [base model card](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) and are **not** measurements of these quantized weights. What you should realistically expect from a quantization of a "think less" fine-tune: token savings largely survive (they come from behaviour, not precision), while the hardest math splits and long-horizon tool chains degrade slightly — most visibly at 4-bit.
|
| 143 |
+
|
| 144 |
+
Measured throughput/acceptance-length data and issue reports (especially "quant X broke task Y") are welcome and will be merged into this table.
|
| 145 |
+
|
| 146 |
+
## Known caveats
|
| 147 |
+
|
| 148 |
+
- Quantization is lossy: expect small regressions versus BF16, largest on competition math and very long agentic traces. Try `oQ6-mtp` before filing a bug.
|
| 149 |
+
- `oQ4-mtp` can amplify repetition on degenerate loops; keep `repetition_penalty` at 1.0 first and only then nudge it.
|
| 150 |
+
- Vision works through the BF16 tower, but I have not benchmarked VQA accuracy post-quantization.
|
| 151 |
+
- 262k context is the architecture's limit, not a promise: keep KV cache within your memory budget or the system swaps.
|
| 152 |
+
- MTP decoding only helps when the speculative draft is enabled in the runtime; without it you pay for the head and get nothing.
|
| 153 |
+
|
| 154 |
+
## License
|
| 155 |
+
|
| 156 |
+
**This repository is distributed under the [Swift Open License v1.0](https://huggingface.co/ukisai/Swift-Qwen3.8-27b#license-and-access).** A quantization is a derivative work: it inherits the upstream terms in full and cannot be released under a more permissive license.
|
| 157 |
+
|
| 158 |
+
- Free personal, research, educational, evaluation and commercial use for individuals and organizations with annual recurring revenue (including affiliates) **up to US$1,000,000**.
|
| 159 |
+
- Above that threshold, commercial use requires a separate **Swift Enterprise License** from UkisAI.
|
| 160 |
+
- The base Qwen3.8 checkpoint and the ThinkingCap-Qwen3.6-27B transfer component (BottleCap AI) contribute their own terms, which apply to you as well — read the `LICENSE` files in [`ukisai/Swift-Qwen3.8-27b`](https://huggingface.co/ukisai/Swift-Qwen3.8-27b) and the BottleCap repository before commercial deployment.
|
| 161 |
+
- Keep this attribution, the upstream citation and the `base_model` metadata intact when you redistribute.
|
| 162 |
+
|
| 163 |
+
## Citation
|
| 164 |
+
|
| 165 |
+
```bibtex
|
| 166 |
+
@misc{swift-qwen3.8-27b-mlx-quants,
|
| 167 |
+
title = {Swift-Qwen3.8-27b-oQ4-mtp}: oMLX/MLX quantization of Swift-Qwen3.8-27B with MTP head retained,
|
| 168 |
+
author = {suzu89},
|
| 169 |
+
year = {2026},
|
| 170 |
+
howpublished = {\url{https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp}},
|
| 171 |
+
note = {Unofficial quantization of ukisai/Swift-Qwen3.8-27b}
|
| 172 |
+
}
|
| 173 |
+
|
| 174 |
+
@misc{swift-qwen3.8-27b,
|
| 175 |
+
title = {Swift-Qwen3.8-27B},
|
| 176 |
+
author = {UkisAI},
|
| 177 |
+
year = {2026},
|
| 178 |
+
url = {https://huggingface.co/ukisai/Swift-Qwen3.8-27b}
|
| 179 |
+
}
|
| 180 |
+
```
|
| 181 |
+
|
| 182 |
+
## Acknowledgements
|
| 183 |
+
|
| 184 |
+
- **UkisAI** for Swift-Qwen3.8-27B and the public evaluation harness.
|
| 185 |
+
- **Qwen team** for the Qwen3.8-27B base model.
|
| 186 |
+
- **BottleCap AI** for the ThinkingCap-Qwen3.6-27B transfer component used upstream.
|
| 187 |
+
- **oMLX** for the Apple Silicon server and quantizer that made these builds possible.
|