File size: 2,274 Bytes
36ff98b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
949fba8
36ff98b
949fba8
36ff98b
949fba8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36ff98b
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
---
library_name: mlx
base_model: ukisai/Swift-Qwen3.8-27b
base_model_relation: quantized
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift-Qwen3.8-27b#license-and-access
pipeline_tag: image-text-to-text
tags:
- mlx
- omlx
- quantized
- mtp
---

# Swift-Qwen3.8-27b-oQ4e-fp16-mtp

## Model architecture

A 27B-class dense multimodal model stored in MLX format. It contains a text backbone, a vision encoder, and one multi-token prediction (MTP) layer.

| Component | Structure |
|---|---|
| Text backbone | 64 layers; hidden size 5,120; feed-forward size 17,408 |
| Attention layout | 48 linear-attention layers and 16 full-attention layers, with full attention every fourth layer |
| Full attention | 24 query heads, 4 key/value heads, head dimension 256 |
| Vocabulary | 248,320 tokens |
| Configured context limit | 262,144 tokens; usable length depends on runtime settings and available memory |
| Vision encoder | 27 layers; hidden size 1,152; 16 attention heads; 16 × 16 image patches |
| Vision-to-text connection | Vision features are projected to the text hidden size of 5,120 |
| MTP | One additional prediction layer with attention and feed-forward projections |

## Weight precision

The oQ4e checkpoint uses mixed precision rather than uniform 4-bit weights:

- The default quantization is **4-bit affine**, with **64 values per group**.
- **187 modules** have explicit **5-bit** overrides in `config.json`.
- The MTP layer's seven large attention and feed-forward matrices use **4-bit** weights.
- The MTP fusion matrix (`mtp.fc`) and normalization weights remain **FP16**. The fusion matrix maps 10,240 input features to 5,120 output features.
- The MTP quantization scales and offsets are stored in **FP16**.

The `fp16` suffix describes the retained floating-point precision; it does not mean the entire model or MTP layer is FP16. Exact per-module settings are recorded in `config.json`.

## Source and license

Source revision: `54e66d6c81439bd4fda5ef9a690fa571e3b0d272`.

Original model by UkisAI. This conversion does not change the upstream Swift Open License v1.0 terms. Consult the [source model license and access information](https://huggingface.co/ukisai/Swift-Qwen3.8-27b#license-and-access).