Yanun commited on
Commit
949fba8
·
verified ·
1 Parent(s): 36ff98b

Replace model description with architecture and precision details

Browse files
Files changed (1) hide show
  1. README.md +24 -3
README.md CHANGED
@@ -15,11 +15,32 @@ tags:
15
 
16
  # Swift-Qwen3.8-27b-oQ4e-fp16-mtp
17
 
18
- Local oMLX oQ4e conversion of [UkisAI Swift-Qwen3.8-27b](https://huggingface.co/ukisai/Swift-Qwen3.8-27b), uploaded by Yanun.
19
 
20
- Converted with oMLX using the oQ4e preset and FP16 floating-point precision, with MTP retained. The `fp16` suffix does not mean that all model or MTP weights are unquantized. See `config.json` for quantization settings and `oq_imatrix_report.json` for calibration results.
21
 
22
- Use a compatible MLX runtime. In oMLX, enable Lightning MTP to use the retained MTP head. Performance depends on hardware, prompts, and runtime settings; no speedup is guaranteed.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
23
 
24
  ## Source and license
25
 
 
15
 
16
  # Swift-Qwen3.8-27b-oQ4e-fp16-mtp
17
 
18
+ ## Model architecture
19
 
20
+ A 27B-class dense multimodal model stored in MLX format. It contains a text backbone, a vision encoder, and one multi-token prediction (MTP) layer.
21
 
22
+ | Component | Structure |
23
+ |---|---|
24
+ | Text backbone | 64 layers; hidden size 5,120; feed-forward size 17,408 |
25
+ | Attention layout | 48 linear-attention layers and 16 full-attention layers, with full attention every fourth layer |
26
+ | Full attention | 24 query heads, 4 key/value heads, head dimension 256 |
27
+ | Vocabulary | 248,320 tokens |
28
+ | Configured context limit | 262,144 tokens; usable length depends on runtime settings and available memory |
29
+ | Vision encoder | 27 layers; hidden size 1,152; 16 attention heads; 16 × 16 image patches |
30
+ | Vision-to-text connection | Vision features are projected to the text hidden size of 5,120 |
31
+ | MTP | One additional prediction layer with attention and feed-forward projections |
32
+
33
+ ## Weight precision
34
+
35
+ The oQ4e checkpoint uses mixed precision rather than uniform 4-bit weights:
36
+
37
+ - The default quantization is **4-bit affine**, with **64 values per group**.
38
+ - **187 modules** have explicit **5-bit** overrides in `config.json`.
39
+ - The MTP layer's seven large attention and feed-forward matrices use **4-bit** weights.
40
+ - The MTP fusion matrix (`mtp.fc`) and normalization weights remain **FP16**. The fusion matrix maps 10,240 input features to 5,120 output features.
41
+ - The MTP quantization scales and offsets are stored in **FP16**.
42
+
43
+ The `fp16` suffix describes the retained floating-point precision; it does not mean the entire model or MTP layer is FP16. Exact per-module settings are recorded in `config.json`.
44
 
45
  ## Source and license
46