TokenAI-zer commited on
Commit
f58e5f7
·
verified ·
1 Parent(s): c1ea340

Update references after account rename

Browse files
Files changed (1) hide show
  1. README.md +5 -5
README.md CHANGED
@@ -35,7 +35,7 @@ I am not affiliated with UkisAI. All upstream weights, benchmarks and license te
35
 
36
  ## Pick a variant
37
 
38
- | | [oQ4-mtp](https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp) | [oQ6-mtp](https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ6-mtp) | [oQ8-mtp](https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ8-mtp) |
39
  |---|---|---|---|
40
  | Weights on disk | 15.81 GiB (16.97 GB), 4 shards | 22.09 GiB (23.72 GB), 5 shards | 27.94 GiB (30.00 GB), 6 shards |
41
  | Weight precision | mixed 4/5-bit | mixed 6/8-bit | uniform 8-bit |
@@ -78,7 +78,7 @@ Precision is mixed per module and recorded verbatim in `config.json` → `quanti
78
 
79
  ```bash
80
  # 1. drop the folder into the oMLX model dir
81
- git clone https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp ~/.omlx/models/Swift-Qwen3.8-27b-oQ4-mtp
82
 
83
  # 2. start the multi-model server (model id = folder name)
84
  omlx serve --model-dir ~/.omlx/models --port 8000
@@ -105,7 +105,7 @@ In oMLX model settings, enable the speculative head and the matching reasoning p
105
 
106
  ```bash
107
  pip install -U mlx-lm mlx-vlm
108
- python -m mlx_lm.server --model suzu89/Swift-Qwen3.8-27b-oQ4-mtp --port 8000
109
  ```
110
 
111
  Only oMLX 0.6.4 is verified by me; if you get `mlx_lm` running this architecture, please open an issue and I will document it.
@@ -162,9 +162,9 @@ Measured throughput/acceptance-length data and issue reports (especially "quant
162
  ```bibtex
163
  @misc{swift-qwen3.8-27b-mlx-quants,
164
  title = {Swift-Qwen3.8-27b-oQ4-mtp}: oMLX/MLX quantization of Swift-Qwen3.8-27B with MTP head retained,
165
- author = {suzu89},
166
  year = {2026},
167
- howpublished = {\url{https://huggingface.co/suzu89/Swift-Qwen3.8-27b-oQ4-mtp}},
168
  note = {Unofficial quantization of ukisai/Swift-Qwen3.8-27b}
169
  }
170
 
 
35
 
36
  ## Pick a variant
37
 
38
+ | | [oQ4-mtp](https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp) | [oQ6-mtp](https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ6-mtp) | [oQ8-mtp](https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ8-mtp) |
39
  |---|---|---|---|
40
  | Weights on disk | 15.81 GiB (16.97 GB), 4 shards | 22.09 GiB (23.72 GB), 5 shards | 27.94 GiB (30.00 GB), 6 shards |
41
  | Weight precision | mixed 4/5-bit | mixed 6/8-bit | uniform 8-bit |
 
78
 
79
  ```bash
80
  # 1. drop the folder into the oMLX model dir
81
+ git clone https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp ~/.omlx/models/Swift-Qwen3.8-27b-oQ4-mtp
82
 
83
  # 2. start the multi-model server (model id = folder name)
84
  omlx serve --model-dir ~/.omlx/models --port 8000
 
105
 
106
  ```bash
107
  pip install -U mlx-lm mlx-vlm
108
+ python -m mlx_lm.server --model TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp --port 8000
109
  ```
110
 
111
  Only oMLX 0.6.4 is verified by me; if you get `mlx_lm` running this architecture, please open an issue and I will document it.
 
162
  ```bibtex
163
  @misc{swift-qwen3.8-27b-mlx-quants,
164
  title = {Swift-Qwen3.8-27b-oQ4-mtp}: oMLX/MLX quantization of Swift-Qwen3.8-27B with MTP head retained,
165
+ author = {TokenAI-zer},
166
  year = {2026},
167
+ howpublished = {\url{https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp}},
168
  note = {Unofficial quantization of ukisai/Swift-Qwen3.8-27b}
169
  }
170