ryugyosoft commited on
Commit
5ddf944
·
verified ·
1 Parent(s): 1adb9f5

Model card in Japanese (README.md) and English (README_en.md); onw 0.4 usage

Browse files
Files changed (2) hide show
  1. README.md +52 -38
  2. README_en.md +68 -0
README.md CHANGED
@@ -11,61 +11,75 @@ tags:
11
  - gemma4
12
  - onw
13
  language:
14
- - en
15
  - ja
 
16
  ---
17
 
18
- # Gemma 4 E4B-it for onw — every component on the Intel NPU
19
 
20
- [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) (text + image) packaged for
21
- **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない), the engine that runs LLMs entirely
22
- on the Intel NPU. Same NPU graphs as the standalone
23
- [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu), run by the shared engine.
24
 
25
- ## Version 2 (compact)
26
 
27
- | | |
28
- |---|---|
29
- | download | 9.2 GB → **4.3 GB** |
30
- | memory in use (working set, NPU 3720) | 16.1 GB → 14.8 GB (without the 64-token block) |
31
- | what changed | decoder weights stored once for the 1 / 16 / 64-token graphs, INT4 per-layer embedding table (host lookup) |
32
- | checked | same answers as v1 on our image chat check (NPU 3720), decode 7.2 tok/s |
33
 
34
- Most of the download saving comes from storing identical weights once; in memory the NPU still keeps one
35
- compiled copy per block size, so the footprint shrinks less. Needs onw 0.3 or newer. The previous version
36
- stays available: `hf download ryugyosoft/gemma-4-E4B-it-onw --revision v1 --local-dir ...`.
 
 
 
 
 
37
 
 
38
 
39
  ```bash
40
  hf download ryugyosoft/onw --local-dir onw && cd onw
41
- start.bat ryugyosoft/gemma-4-E4B-it-onw # Windows
42
- bash start.sh ryugyosoft/gemma-4-E4B-it-onw # Ubuntu
43
  ```
44
 
45
- | NPU 3720 (Core Ultra 9 285HX) | |
46
- |---|---|
47
- | decode | **6.7-6.8 tok/s** (prompt lookup decoding on, exact) |
48
- | image + question (284 tokens) | vision 1.9 s + prefill 2.0 s (64-token blocks, >= 24 GB RAM) |
49
- | follow-up turn | only new tokens are processed (turn 2 of an image chat: 1.7 s vs 4.8 s) |
50
- | memory | ~8 GB with the 64-token block, ~6 GB without |
 
 
51
 
52
- What changed vs the standalone repo: the token embedding and the 2.7 GB per-layer embedding table are looked up
53
- on the host (they are table reads; no NPU worker process any more), the LM head is one shared INT8 input for all
54
- block sizes, and prompt lookup decoding / prefix reuse come from the engine.
55
 
56
- **Download with `hf download` or let `onw` fetch it** — a `git clone` without Git LFS gets pointer files, which
57
- `onw` detects and reports.
58
 
59
- ## Files
60
 
61
- | file | what |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
  |---|---|
63
- | `seg00_S{1,16,64}.xml` + `seg00.bin` | the static Gemma 4 decoder (head_dim-512 attention split into 256-wide heads, host-built masks) for 1 / 16 / 64-token blocks |
64
- | `seg01_S*.xml` | LM head + logit softcapping (shared INT8 weight) |
65
- | `vision.xml` | static vision encoder (280 soft tokens) |
66
- | `shared.bin` | INT8 token embedding (= tied head), INT4 per-layer embedding table + id map |
67
- | `engine.json`, tokenizer / processor files | metadata for onw |
 
 
 
 
68
 
69
- ## License
70
 
71
- Apache 2.0, same as the base model. The weights are re-quantized / restructured from it.
 
11
  - gemma4
12
  - onw
13
  language:
 
14
  - ja
15
+ - en
16
  ---
17
 
18
+ # gemma-4-E4B-it for onw — Intel NPU だけで動く(テキスト+画像)
19
 
20
+ **日本語** | [English](README_en.md)
 
 
 
21
 
22
+ [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) を **[onw](https://huggingface.co/ryugyosoft/onw)**(俺のNPUがこんなに動くわけない)用に変換したモデルです。onw は LLM を Intel NPU だけで動かすエンジンで、このリポジトリにはモデルだけが入っています(エンジンは別にダウンロードします)。
23
 
24
+ テキストと画像に対応。単独版 [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu) と同じ検証済みの NPU グラフを、共通エンジンで動かします。
 
 
 
 
 
25
 
26
+ | NPU 3720(Core Ultra 9 285HX) | |
27
+ |---|---|
28
+ | 入力 | テキスト+画像 |
29
+ | 生成速度 | **7.2 tok/s(先読み検証あり、出力は通常生成と同一)** |
30
+ | プロンプト処理 | 画像+質問 284 トークン:画像 1.9 秒+処理 2.0 秒(64 トークン版、RAM 24 GB 以上) |
31
+ | ダウンロード | **4.3 GB** |
32
+ | 使用メモリ(読み込み後の作業セット) | 約 15 GB(64 トークン版なし) |
33
+ | 初回起動(NPU 向けコンパイル)/2 回目以降 | 約 4 分/約 20 秒 |
34
 
35
+ ## 使い方
36
 
37
  ```bash
38
  hf download ryugyosoft/onw --local-dir onw && cd onw
39
+ setup.bat # Windows(Ubuntu は bash setup.sh)
 
40
  ```
41
 
42
+ 常駐アプリが起動し、onw ウィンドウの「モデル」タブが開きます。**gemma-4-E4B-it** を選んで「ダウンロード」を押し、「サーバー」タブで起動してください。起動後は OpenAI 互換 API として他のアプリから使えます。
43
+
44
+ ```python
45
+ from openai import OpenAI
46
+ c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
47
+ r = c.chat.completions.create(model="gemma-4-E4B-it-onw", messages=[{"role": "user", "content": "こんにちは"}])
48
+ print(r.choices[0].message.content)
49
+ ```
50
 
51
+ コンソールで一度だけ起動する場合は `start.bat ryugyosoft/gemma-4-E4B-it-onw`(Ubuntu は `bash start.sh ryugyosoft/gemma-4-E4B-it-onw`)。動作確認用のチャット画面は `/chat`、速度などを一括で確かめる検証ページは `/check` です。
 
 
52
 
53
+ **`git clone` ではなく `hf download`(または onw からのダウンロード)を使ってください。** Git LFS なしで clone すると重みの代わりに小さなポインタファイルしか入らず、onw はそれを検出して止まります。
 
54
 
55
+ ## バージョン
56
 
57
+ - **v2(現在)**:デコーダーの重みを 1/16/64 トークン用で共有、層ごとの埋め込み表を INT4 化(9.2 → 4.3 GB、使用メモリ 16.1 → 14.8 GB)。回答は v1 と同じでした。 onw 0.3 以降が必要です。
58
+ - v1:`hf download ryugyosoft/gemma-4-E4B-it-onw --revision v1 --local-dir ...` で取り出せます。
59
+
60
+ ダウンロード量の削減の大半は同じ重みを 1 つにまとめたことによるもので、実行時は NPU がブロックサイズごとにコンパイル済みの重みを持つため、使用メモリの減り方は小さめです。
61
+
62
+ ## 仕組み
63
+
64
+ - Gemma 4 用に書き換えた静的デコーダー(ヘッド次元 512 のアテンションを 256 幅に分割、マスクはホストで作成)をそのまま取り込んでいます。
65
+ - 埋め込みと 2.7 GB の層ごとの埋め込み表はホストで引きます(表を引くだけなので NPU メモリを使わず、単独版にあった別プロセスも不要)。
66
+ - 出力層は全ブロックサイズで共有する INT8 入力。logit の softcap は NPU コンパイラの不具合を避ける形で組んでいます。
67
+ - 詳しくは [onw の README](https://huggingface.co/ryugyosoft/onw)。
68
+
69
+ ## ファイル
70
+
71
+ | ファイル | 内容 |
72
  |---|---|
73
+ | `seg00_S{1,16,64}.xml` + `seg00.bin` | 静的デコーダー(1/16/64 トークン用、重みは共有) |
74
+ | `seg01_S*.xml` + `seg01.bin` | 出力層と logit softcap |
75
+ | `vision.xml` | 静的な画像エンコーダー(280 トークン) |
76
+ | `shared.bin` | INT8 の埋め込み(出力層と共有)、INT4 の層ごとの埋め込み表と ID 変換表 |
77
+ | `engine.json`、トークナイザー/設定ファイル | onw 用のメタデータ、元のリポジトリの設定 |
78
+
79
+ ## 品質
80
+
81
+ エキスパートなどはチャネル単位・グループ単位の INT4(四捨五入量子化)です。確認した範囲では最初のトークンの予測は bf16 と一致し、回答は自然ですが、数トークン先から言い回しが元のモデルと分かれることがあります。
82
 
83
+ ## ライセンス
84
 
85
+ Apache 2.0(元モデルと同じ)。重みは元モデルから再量子化・再構成したものです。
README_en.md ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # gemma-4-E4B-it for onw — entirely on the Intel NPU (text + image)
2
+
3
+ [日本語](README.md) | **English**
4
+
5
+ [google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) converted for **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.
6
+
7
+ Text + image. The same validated NPU graphs as the standalone [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu), run by the shared engine.
8
+
9
+ | NPU 3720 (Core Ultra 9 285HX) | |
10
+ |---|---|
11
+ | input | text + image |
12
+ | decode | **7.2 tok/s (prompt lookup on, same output as plain greedy)** |
13
+ | prompt processing | image + question, 284 tokens: vision 1.9 s + prefill 2.0 s (64-token blocks, >= 24 GB RAM) |
14
+ | download | **4.3 GB** |
15
+ | memory in use (working set after loading) | ~15 GB (without the 64-token block) |
16
+ | first start (NPU compile) / later | ~4 min / ~20 s |
17
+
18
+ ## Use
19
+
20
+ ```bash
21
+ hf download ryugyosoft/onw --local-dir onw && cd onw
22
+ setup.bat # Windows (Ubuntu: bash setup.sh)
23
+ ```
24
+
25
+ The resident app starts and the onw window opens on its Models tab: pick **gemma-4-E4B-it**, press Download, then start it on the Server tab. It is then an OpenAI-compatible API for other apps:
26
+
27
+ ```python
28
+ from openai import OpenAI
29
+ c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
30
+ r = c.chat.completions.create(model="gemma-4-E4B-it-onw", messages=[{"role": "user", "content": "Hello"}])
31
+ print(r.choices[0].message.content)
32
+ ```
33
+
34
+ To run it once from a console: `start.bat ryugyosoft/gemma-4-E4B-it-onw` (Ubuntu: `bash start.sh ryugyosoft/gemma-4-E4B-it-onw`). A chat UI for trying it is at `/chat`, a check page with speed tests at `/check`.
35
+
36
+ **Download with `hf download` (or from onw), not `git clone`.** A clone without Git LFS gets small pointer files instead of the weights; onw detects that and stops.
37
+
38
+ ## Versions
39
+
40
+ - **v2 (current)**: decoder weights stored once for the 1 / 16 / 64-token graphs, INT4 per-layer embedding table (9.2 → 4.3 GB; memory 16.1 → 14.8 GB). Same answers as v1 on our checks. Needs onw 0.3 or newer.
41
+ - v1: `hf download ryugyosoft/gemma-4-E4B-it-onw --revision v1 --local-dir ...`
42
+
43
+ Most of the download saving comes from storing identical weights once; in memory the NPU keeps one compiled copy per block size, so the footprint shrinks less.
44
+
45
+ ## How it runs
46
+
47
+ - The static Gemma 4 decoder (head_dim-512 attention split into 256-wide heads, host-built masks) is imported as is.
48
+ - The token embedding and the 2.7 GB per-layer embedding table are looked up on the host (table reads: no NPU memory, no worker process).
49
+ - The LM head is one shared INT8 input for all block sizes; the logit softcap is arranged to avoid an NPU compiler bug.
50
+ - Details: [onw README](https://huggingface.co/ryugyosoft/onw).
51
+
52
+ ## Files
53
+
54
+ | file | what |
55
+ |---|---|
56
+ | `seg00_S{1,16,64}.xml` + `seg00.bin` | static decoder for 1 / 16 / 64-token blocks (shared weights) |
57
+ | `seg01_S*.xml` + `seg01.bin` | LM head + logit softcapping |
58
+ | `vision.xml` | static vision encoder (280 soft tokens) |
59
+ | `shared.bin` | INT8 token embedding (= tied head), INT4 per-layer embedding table + id map |
60
+ | `engine.json`, tokenizer / config files | onw metadata, the original repo's configs |
61
+
62
+ ## Quality
63
+
64
+ Channel-wise / group-wise INT4 (round-to-nearest): on our checks the first token matches the bf16 model; answers are fluent, word choices may diverge after a few tokens.
65
+
66
+ ## License
67
+
68
+ Apache 2.0, same as the base model. The weights are re-quantized / restructured from it.