ryugyosoft commited on
Commit
1adb9f5
·
verified ·
1 Parent(s): 0c1184e

v2: compact (4.3 GB)

Browse files
README.md CHANGED
@@ -1,57 +1,71 @@
1
- ---
2
- license: apache-2.0
3
- base_model: google/gemma-4-E4B-it
4
- pipeline_tag: image-text-to-text
5
- library_name: openvino
6
- tags:
7
- - openvino
8
- - npu
9
- - intel
10
- - intel-npu
11
- - gemma4
12
- - onw
13
- language:
14
- - en
15
- - ja
16
- ---
17
-
18
- # Gemma 4 E4B-it for onw — every component on the Intel NPU
19
-
20
- [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) (text + image) packaged for
21
- **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない), the engine that runs LLMs entirely
22
- on the Intel NPU. Same NPU graphs as the standalone
23
- [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu), run by the shared engine.
24
-
25
- ```bash
26
- hf download ryugyosoft/onw --local-dir onw && cd onw
27
- start.bat ryugyosoft/gemma-4-E4B-it-onw # Windows
28
- bash start.sh ryugyosoft/gemma-4-E4B-it-onw # Ubuntu
29
- ```
30
-
31
- | NPU 3720 (Core Ultra 9 285HX) | |
32
- |---|---|
33
- | decode | **6.7-6.8 tok/s** (prompt lookup decoding on, exact) |
34
- | image + question (284 tokens) | vision 1.9 s + prefill 2.0 s (64-token blocks, >= 24 GB RAM) |
35
- | follow-up turn | only new tokens are processed (turn 2 of an image chat: 1.7 s vs 4.8 s) |
36
- | memory | ~8 GB with the 64-token block, ~6 GB without |
37
-
38
- What changed vs the standalone repo: the token embedding and the 2.7 GB per-layer embedding table are looked up
39
- on the host (they are table reads; no NPU worker process any more), the LM head is one shared INT8 input for all
40
- block sizes, and prompt lookup decoding / prefix reuse come from the engine.
41
-
42
- **Download with `hf download` or let `onw` fetch it** — a `git clone` without Git LFS gets pointer files, which
43
- `onw` detects and reports.
44
-
45
- ## Files
46
-
47
- | file | what |
48
- |---|---|
49
- | `seg00_S{1,16,64}.xml` | the static Gemma 4 decoder (head_dim-512 attention split into 256-wide heads, host-built masks) for 1 / 16 / 64-token blocks |
50
- | `seg01_S*.xml` | LM head + logit softcapping (shared INT8 weight) |
51
- | `vision.xml` | static vision encoder (280 soft tokens) |
52
- | `shared.bin` | INT8 token embedding (= tied head), per-layer embedding table + id map |
53
- | `engine.json`, tokenizer / processor files | metadata for onw |
54
-
55
- ## License
56
-
57
- Apache 2.0, same as the base model. The weights are re-quantized / restructured from it.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: google/gemma-4-E4B-it
4
+ pipeline_tag: image-text-to-text
5
+ library_name: openvino
6
+ tags:
7
+ - openvino
8
+ - npu
9
+ - intel
10
+ - intel-npu
11
+ - gemma4
12
+ - onw
13
+ language:
14
+ - en
15
+ - ja
16
+ ---
17
+
18
+ # Gemma 4 E4B-it for onw — every component on the Intel NPU
19
+
20
+ [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) (text + image) packaged for
21
+ **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない), the engine that runs LLMs entirely
22
+ on the Intel NPU. Same NPU graphs as the standalone
23
+ [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu), run by the shared engine.
24
+
25
+ ## Version 2 (compact)
26
+
27
+ | | |
28
+ |---|---|
29
+ | download | 9.2 GB → **4.3 GB** |
30
+ | memory in use (working set, NPU 3720) | 16.1 GB → 14.8 GB (without the 64-token block) |
31
+ | what changed | decoder weights stored once for the 1 / 16 / 64-token graphs, INT4 per-layer embedding table (host lookup) |
32
+ | checked | same answers as v1 on our image chat check (NPU 3720), decode 7.2 tok/s |
33
+
34
+ Most of the download saving comes from storing identical weights once; in memory the NPU still keeps one
35
+ compiled copy per block size, so the footprint shrinks less. Needs onw 0.3 or newer. The previous version
36
+ stays available: `hf download ryugyosoft/gemma-4-E4B-it-onw --revision v1 --local-dir ...`.
37
+
38
+
39
+ ```bash
40
+ hf download ryugyosoft/onw --local-dir onw && cd onw
41
+ start.bat ryugyosoft/gemma-4-E4B-it-onw # Windows
42
+ bash start.sh ryugyosoft/gemma-4-E4B-it-onw # Ubuntu
43
+ ```
44
+
45
+ | NPU 3720 (Core Ultra 9 285HX) | |
46
+ |---|---|
47
+ | decode | **6.7-6.8 tok/s** (prompt lookup decoding on, exact) |
48
+ | image + question (284 tokens) | vision 1.9 s + prefill 2.0 s (64-token blocks, >= 24 GB RAM) |
49
+ | follow-up turn | only new tokens are processed (turn 2 of an image chat: 1.7 s vs 4.8 s) |
50
+ | memory | ~8 GB with the 64-token block, ~6 GB without |
51
+
52
+ What changed vs the standalone repo: the token embedding and the 2.7 GB per-layer embedding table are looked up
53
+ on the host (they are table reads; no NPU worker process any more), the LM head is one shared INT8 input for all
54
+ block sizes, and prompt lookup decoding / prefix reuse come from the engine.
55
+
56
+ **Download with `hf download` or let `onw` fetch it** — a `git clone` without Git LFS gets pointer files, which
57
+ `onw` detects and reports.
58
+
59
+ ## Files
60
+
61
+ | file | what |
62
+ |---|---|
63
+ | `seg00_S{1,16,64}.xml` + `seg00.bin` | the static Gemma 4 decoder (head_dim-512 attention split into 256-wide heads, host-built masks) for 1 / 16 / 64-token blocks |
64
+ | `seg01_S*.xml` | LM head + logit softcapping (shared INT8 weight) |
65
+ | `vision.xml` | static vision encoder (280 soft tokens) |
66
+ | `shared.bin` | INT8 token embedding (= tied head), INT4 per-layer embedding table + id map |
67
+ | `engine.json`, tokenizer / processor files | metadata for onw |
68
+
69
+ ## License
70
+
71
+ Apache 2.0, same as the base model. The weights are re-quantized / restructured from it.
engine.json CHANGED
@@ -105,7 +105,8 @@
105
  "router_layer": null,
106
  "S": 1,
107
  "slots": 0,
108
- "head": false
 
109
  },
110
  "seg01_S1.xml": {
111
  "index": 1,
@@ -113,7 +114,8 @@
113
  "router_layer": null,
114
  "S": 1,
115
  "slots": 0,
116
- "head": true
 
117
  },
118
  "seg00_S16.xml": {
119
  "index": 0,
@@ -121,7 +123,8 @@
121
  "router_layer": null,
122
  "S": 16,
123
  "slots": 0,
124
- "head": false
 
125
  },
126
  "seg01_S16.xml": {
127
  "index": 1,
@@ -129,7 +132,8 @@
129
  "router_layer": null,
130
  "S": 16,
131
  "slots": 0,
132
- "head": true
 
133
  },
134
  "seg00_S64.xml": {
135
  "index": 0,
@@ -137,7 +141,8 @@
137
  "router_layer": null,
138
  "S": 64,
139
  "slots": 0,
140
- "head": false
 
141
  },
142
  "seg01_S64.xml": {
143
  "index": 1,
@@ -145,7 +150,8 @@
145
  "router_layer": null,
146
  "S": 64,
147
  "slots": 0,
148
- "head": true
 
149
  }
150
  },
151
  "masks": {
@@ -193,5 +199,11 @@
193
  "NPU_USE_NPUW": "YES",
194
  "NPUW_DEVICES": "NPU",
195
  "NPUW_FOLD": "YES"
 
 
 
 
 
 
196
  }
197
  }
 
105
  "router_layer": null,
106
  "S": 1,
107
  "slots": 0,
108
+ "head": false,
109
+ "bin": "seg00.bin"
110
  },
111
  "seg01_S1.xml": {
112
  "index": 1,
 
114
  "router_layer": null,
115
  "S": 1,
116
  "slots": 0,
117
+ "head": true,
118
+ "bin": "seg01.bin"
119
  },
120
  "seg00_S16.xml": {
121
  "index": 0,
 
123
  "router_layer": null,
124
  "S": 16,
125
  "slots": 0,
126
+ "head": false,
127
+ "bin": "seg00.bin"
128
  },
129
  "seg01_S16.xml": {
130
  "index": 1,
 
132
  "router_layer": null,
133
  "S": 16,
134
  "slots": 0,
135
+ "head": true,
136
+ "bin": "seg01.bin"
137
  },
138
  "seg00_S64.xml": {
139
  "index": 0,
 
141
  "router_layer": null,
142
  "S": 64,
143
  "slots": 0,
144
+ "head": false,
145
+ "bin": "seg00.bin"
146
  },
147
  "seg01_S64.xml": {
148
  "index": 1,
 
150
  "router_layer": null,
151
  "S": 64,
152
  "slots": 0,
153
+ "head": true,
154
+ "bin": "seg01.bin"
155
  }
156
  },
157
  "masks": {
 
199
  "NPU_USE_NPUW": "YES",
200
  "NPUW_DEVICES": "NPU",
201
  "NPUW_FOLD": "YES"
202
+ },
203
+ "tables": {
204
+ "ple": {
205
+ "bits": 4,
206
+ "group": 32
207
+ }
208
  }
209
  }
seg00_S1.bin → seg00.bin RENAMED
File without changes
seg00_S16.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:38b18a93e134fd55cdcd54772c26b4d966a371c6763eb8ab020e0ce91f87a971
3
- size 2050670280
 
 
 
 
seg00_S64.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:38b18a93e134fd55cdcd54772c26b4d966a371c6763eb8ab020e0ce91f87a971
3
- size 2050670280
 
 
 
 
seg01_S1.bin → seg01.bin RENAMED
File without changes
seg01_S16.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:5e48af94438b6e5de715ec5e3e80e252dc3f66d9ea35b3bdb8dac0650fa8ad71
3
- size 6
 
 
 
 
seg01_S64.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:5e48af94438b6e5de715ec5e3e80e252dc3f66d9ea35b3bdb8dac0650fa8ad71
3
- size 6
 
 
 
 
shared.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:402bac867a19a87a5b21800b3a093bc56306b49e52a7bd5a3a9cdf471a7175c6
3
- size 3492282368
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9800fc61111c52fff3c69497b1cff3e2623af746ae55a38c60733580c57c3e3a
3
+ size 2258632704
shared.bin.json CHANGED
@@ -1 +1 @@
1
- {"tensors": {"embed_q": {"at": [0, 671088640], "type": "i8", "shape": [262144, 2560]}, "embed_s": {"at": [671088640, 524288], "type": "f16", "shape": [262144]}, "head_q": {"alias": "embed_q"}, "head_s": {"at": [671612928, 524288], "type": "f16", "shape": [262144]}, "ple_q": {"at": [672137216, 2818572288], "type": "i8", "shape": [262144, 10752]}, "ple_s": {"at": [3490709504, 524288], "type": "f16", "shape": [262144]}, "ple_map": {"at": [3491233792, 1048576], "type": "i32", "shape": [262144]}}}
 
1
+ {"tensors": {"ple_q4": {"at": [0, 1409286144], "type": "u8", "shape": [262144, 5376]}, "ple_s4": {"at": [1409286144, 176160768], "type": "f16", "shape": [262144, 336]}, "embed_q": {"at": [1585446912, 671088640], "type": "i8", "shape": [262144, 2560]}, "embed_s": {"at": [2256535552, 524288], "type": "f16", "shape": [262144]}, "head_q": {"alias": "embed_q"}, "head_s": {"at": [2257059840, 524288], "type": "f16", "shape": [262144]}, "ple_map": {"at": [2257584128, 1048576], "type": "i32", "shape": [262144]}}}