aufklarer commited on
Commit
2956b1b
·
verified ·
1 Parent(s): 0fd741d

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +61 -0
README.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-0.8B
4
+ tags:
5
+ - mlx
6
+ - qwen3.5
7
+ - text-generation
8
+ - on-device
9
+ - apple-silicon
10
+ - int4
11
+ - int8
12
+ language:
13
+ - en
14
+ - zh
15
+ - ja
16
+ - ko
17
+ - de
18
+ - es
19
+ - fr
20
+ library_name: mlx
21
+ pipeline_tag: text-generation
22
+ ---
23
+
24
+ # Qwen3.5-0.8B Chat — MLX (Apple Silicon)
25
+
26
+ Text-only extraction of [Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) for on-device LLM chat on Apple Silicon via [MLX](https://github.com/ml-explore/mlx).
27
+
28
+ ## Architecture
29
+
30
+ Qwen3.5 is a **hybrid** model with 24 layers:
31
+ - **18× DeltaNet** — linear attention with gated delta rule recurrence, O(1) memory per step
32
+ - **6× GatedAttention** — full scaled dot-product attention with KV cache, partial RoPE (25%)
33
+ - Pattern: `[linear, linear, linear, full] × 6`
34
+ - Tied word embeddings (lm_head = embed_tokens)
35
+
36
+ ## Variants
37
+
38
+ | Variant | Size | Path |
39
+ |---------|------|------|
40
+ | INT4 | 404 MB | `int4/model.safetensors` |
41
+ | INT8 | 763 MB | `int8/model.safetensors` |
42
+
43
+ Each variant includes `config.json`, `tokenizer.json`, and `tokenizer_config.json`.
44
+
45
+ ## Usage
46
+
47
+ ```swift
48
+ import Qwen3Chat
49
+
50
+ let model = try await Qwen35MLXChat.fromPretrained(quantization: .int4)
51
+ let response = try model.generate(
52
+ messages: [ChatMessage(role: .user, content: "Hello!")],
53
+ sampling: ChatSamplingConfig(temperature: 0.3, maxTokens: 100)
54
+ )
55
+ ```
56
+
57
+ Part of the [soniqo](https://soniqo.audio) speech toolkit for Apple Silicon.
58
+
59
+ ## Source
60
+
61
+ Repackaged from [mlx-community/Qwen3.5-0.8B-4bit](https://huggingface.co/mlx-community/Qwen3.5-0.8B-4bit) and [mlx-community/Qwen3.5-0.8B-MLX-8bit](https://huggingface.co/mlx-community/Qwen3.5-0.8B-MLX-8bit) — vision tower removed, text model only.