Model card in Japanese (README.md) and English (README_en.md); onw 0.4 usage
Browse files- README.md +52 -38
- README_en.md +68 -0
README.md
CHANGED
|
@@ -11,61 +11,75 @@ tags:
|
|
| 11 |
- gemma4
|
| 12 |
- onw
|
| 13 |
language:
|
| 14 |
-
- en
|
| 15 |
- ja
|
|
|
|
| 16 |
---
|
| 17 |
|
| 18 |
-
#
|
| 19 |
|
| 20 |
-
[
|
| 21 |
-
**[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない), the engine that runs LLMs entirely
|
| 22 |
-
on the Intel NPU. Same NPU graphs as the standalone
|
| 23 |
-
[ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu), run by the shared engine.
|
| 24 |
|
| 25 |
-
|
| 26 |
|
| 27 |
-
|
| 28 |
-
|---|---|
|
| 29 |
-
| download | 9.2 GB → **4.3 GB** |
|
| 30 |
-
| memory in use (working set, NPU 3720) | 16.1 GB → 14.8 GB (without the 64-token block) |
|
| 31 |
-
| what changed | decoder weights stored once for the 1 / 16 / 64-token graphs, INT4 per-layer embedding table (host lookup) |
|
| 32 |
-
| checked | same answers as v1 on our image chat check (NPU 3720), decode 7.2 tok/s |
|
| 33 |
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
|
|
|
| 38 |
|
| 39 |
```bash
|
| 40 |
hf download ryugyosoft/onw --local-dir onw && cd onw
|
| 41 |
-
|
| 42 |
-
bash start.sh ryugyosoft/gemma-4-E4B-it-onw # Ubuntu
|
| 43 |
```
|
| 44 |
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
|
|
|
|
|
|
| 51 |
|
| 52 |
-
|
| 53 |
-
on the host (they are table reads; no NPU worker process any more), the LM head is one shared INT8 input for all
|
| 54 |
-
block sizes, and prompt lookup decoding / prefix reuse come from the engine.
|
| 55 |
|
| 56 |
-
**
|
| 57 |
-
`onw` detects and reports.
|
| 58 |
|
| 59 |
-
##
|
| 60 |
|
| 61 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|---|---|
|
| 63 |
-
| `seg00_S{1,16,64}.xml` + `seg00.bin` |
|
| 64 |
-
| `seg01_S*.xml`
|
| 65 |
-
| `vision.xml` |
|
| 66 |
-
| `shared.bin` | INT8
|
| 67 |
-
| `engine.json`
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
-
##
|
| 70 |
|
| 71 |
-
Apache 2.0
|
|
|
|
| 11 |
- gemma4
|
| 12 |
- onw
|
| 13 |
language:
|
|
|
|
| 14 |
- ja
|
| 15 |
+
- en
|
| 16 |
---
|
| 17 |
|
| 18 |
+
# gemma-4-E4B-it for onw — Intel NPU だけで動く(テキスト+画像)
|
| 19 |
|
| 20 |
+
**日本語** | [English](README_en.md)
|
|
|
|
|
|
|
|
|
|
| 21 |
|
| 22 |
+
[google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) を **[onw](https://huggingface.co/ryugyosoft/onw)**(俺のNPUがこんなに動くわけない)用に変換したモデルです。onw は LLM を Intel NPU だけで動かすエンジンで、このリポジトリにはモデルだけが入っています(エンジンは別にダウンロードします)。
|
| 23 |
|
| 24 |
+
テキストと画像に対応。単独版 [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu) と同じ検証済みの NPU グラフを、共通エンジンで動かします。
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
+
| NPU 3720(Core Ultra 9 285HX) | |
|
| 27 |
+
|---|---|
|
| 28 |
+
| 入力 | テキスト+画像 |
|
| 29 |
+
| 生成速度 | **7.2 tok/s(先読み検証あり、出力は通常生成と同一)** |
|
| 30 |
+
| プロンプト処理 | 画像+質問 284 トークン:画像 1.9 秒+処理 2.0 秒(64 トークン版、RAM 24 GB 以上) |
|
| 31 |
+
| ダウンロード | **4.3 GB** |
|
| 32 |
+
| 使用メモリ(読み込み後の作業セット) | 約 15 GB(64 トークン版なし) |
|
| 33 |
+
| 初回起動(NPU 向けコンパイル)/2 回目以降 | 約 4 分/約 20 秒 |
|
| 34 |
|
| 35 |
+
## 使い方
|
| 36 |
|
| 37 |
```bash
|
| 38 |
hf download ryugyosoft/onw --local-dir onw && cd onw
|
| 39 |
+
setup.bat # Windows(Ubuntu は bash setup.sh)
|
|
|
|
| 40 |
```
|
| 41 |
|
| 42 |
+
常駐アプリが起動し、onw ウィンドウの「モデル」タブが開きます。**gemma-4-E4B-it** を選んで「ダウンロード」を押し、「サーバー」タブで起動してください。起動後は OpenAI 互換 API として他のアプリから使えます。
|
| 43 |
+
|
| 44 |
+
```python
|
| 45 |
+
from openai import OpenAI
|
| 46 |
+
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
|
| 47 |
+
r = c.chat.completions.create(model="gemma-4-E4B-it-onw", messages=[{"role": "user", "content": "こんにちは"}])
|
| 48 |
+
print(r.choices[0].message.content)
|
| 49 |
+
```
|
| 50 |
|
| 51 |
+
コンソールで一度だけ起動する場合は `start.bat ryugyosoft/gemma-4-E4B-it-onw`(Ubuntu は `bash start.sh ryugyosoft/gemma-4-E4B-it-onw`)。動作確認用のチャット画面は `/chat`、速度などを一括で確かめる検証ページは `/check` です。
|
|
|
|
|
|
|
| 52 |
|
| 53 |
+
**`git clone` ではなく `hf download`(または onw からのダウンロード)を使ってください。** Git LFS なしで clone すると重みの代わりに小さなポインタファイルしか入らず、onw はそれを検出して止まります。
|
|
|
|
| 54 |
|
| 55 |
+
## バージョン
|
| 56 |
|
| 57 |
+
- **v2(現在)**:デコーダーの重みを 1/16/64 トークン用で共有、層ごとの埋め込み表を INT4 化(9.2 → 4.3 GB、使用メモリ 16.1 → 14.8 GB)。回答は v1 と同じでした。 onw 0.3 以降が必要です。
|
| 58 |
+
- v1:`hf download ryugyosoft/gemma-4-E4B-it-onw --revision v1 --local-dir ...` で取り出せます。
|
| 59 |
+
|
| 60 |
+
ダウンロード量の削減の大半は同じ重みを 1 つにまとめたことによるもので、実行時は NPU がブロックサイズごとにコンパイル済みの重みを持つため、使用メモリの減り方は小さめです。
|
| 61 |
+
|
| 62 |
+
## 仕組み
|
| 63 |
+
|
| 64 |
+
- Gemma 4 用に書き換えた静的デコーダー(ヘッド次元 512 のアテンションを 256 幅に分割、マスクはホストで作成)をそのまま取り込んでいます。
|
| 65 |
+
- 埋め込みと 2.7 GB の層ごとの埋め込み表はホストで引きます(表を引くだけなので NPU メモリを使わず、単独版にあった別プロセスも不要)。
|
| 66 |
+
- 出力層は全ブロックサイズで共有する INT8 入力。logit の softcap は NPU コンパイラの不具合を避ける形で組んでいます。
|
| 67 |
+
- 詳しくは [onw の README](https://huggingface.co/ryugyosoft/onw)。
|
| 68 |
+
|
| 69 |
+
## ファイル
|
| 70 |
+
|
| 71 |
+
| ファイル | 内容 |
|
| 72 |
|---|---|
|
| 73 |
+
| `seg00_S{1,16,64}.xml` + `seg00.bin` | 静的デコーダー(1/16/64 トークン用、重みは共有) |
|
| 74 |
+
| `seg01_S*.xml` + `seg01.bin` | 出力層と logit softcap |
|
| 75 |
+
| `vision.xml` | 静的な画像エンコーダー(280 トークン) |
|
| 76 |
+
| `shared.bin` | INT8 の埋め込み(出力層と共有)、INT4 の層ごとの埋め込み表と ID 変換表 |
|
| 77 |
+
| `engine.json`、トークナイザー/設定ファイル | onw 用のメタデータ、元のリポジトリの設定 |
|
| 78 |
+
|
| 79 |
+
## 品質
|
| 80 |
+
|
| 81 |
+
エキスパートなどはチャネル単位・グループ単位の INT4(四捨五入量子化)です。確認した範囲では最初のトークンの予測は bf16 と一致し、回答は自然ですが、数トークン先から言い回しが元のモデルと分かれることがあります。
|
| 82 |
|
| 83 |
+
## ライセンス
|
| 84 |
|
| 85 |
+
Apache 2.0(元モデルと同じ)。重みは元モデルから再量子化・再構成したものです。
|
README_en.md
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# gemma-4-E4B-it for onw — entirely on the Intel NPU (text + image)
|
| 2 |
+
|
| 3 |
+
[日本語](README.md) | **English**
|
| 4 |
+
|
| 5 |
+
[google/gemma-4-E4B-it](https://huggingface.co/google/gemma-4-E4B-it) converted for **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.
|
| 6 |
+
|
| 7 |
+
Text + image. The same validated NPU graphs as the standalone [ryugyosoft/gemma-4-E4B-it-npu](https://huggingface.co/ryugyosoft/gemma-4-E4B-it-npu), run by the shared engine.
|
| 8 |
+
|
| 9 |
+
| NPU 3720 (Core Ultra 9 285HX) | |
|
| 10 |
+
|---|---|
|
| 11 |
+
| input | text + image |
|
| 12 |
+
| decode | **7.2 tok/s (prompt lookup on, same output as plain greedy)** |
|
| 13 |
+
| prompt processing | image + question, 284 tokens: vision 1.9 s + prefill 2.0 s (64-token blocks, >= 24 GB RAM) |
|
| 14 |
+
| download | **4.3 GB** |
|
| 15 |
+
| memory in use (working set after loading) | ~15 GB (without the 64-token block) |
|
| 16 |
+
| first start (NPU compile) / later | ~4 min / ~20 s |
|
| 17 |
+
|
| 18 |
+
## Use
|
| 19 |
+
|
| 20 |
+
```bash
|
| 21 |
+
hf download ryugyosoft/onw --local-dir onw && cd onw
|
| 22 |
+
setup.bat # Windows (Ubuntu: bash setup.sh)
|
| 23 |
+
```
|
| 24 |
+
|
| 25 |
+
The resident app starts and the onw window opens on its Models tab: pick **gemma-4-E4B-it**, press Download, then start it on the Server tab. It is then an OpenAI-compatible API for other apps:
|
| 26 |
+
|
| 27 |
+
```python
|
| 28 |
+
from openai import OpenAI
|
| 29 |
+
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
|
| 30 |
+
r = c.chat.completions.create(model="gemma-4-E4B-it-onw", messages=[{"role": "user", "content": "Hello"}])
|
| 31 |
+
print(r.choices[0].message.content)
|
| 32 |
+
```
|
| 33 |
+
|
| 34 |
+
To run it once from a console: `start.bat ryugyosoft/gemma-4-E4B-it-onw` (Ubuntu: `bash start.sh ryugyosoft/gemma-4-E4B-it-onw`). A chat UI for trying it is at `/chat`, a check page with speed tests at `/check`.
|
| 35 |
+
|
| 36 |
+
**Download with `hf download` (or from onw), not `git clone`.** A clone without Git LFS gets small pointer files instead of the weights; onw detects that and stops.
|
| 37 |
+
|
| 38 |
+
## Versions
|
| 39 |
+
|
| 40 |
+
- **v2 (current)**: decoder weights stored once for the 1 / 16 / 64-token graphs, INT4 per-layer embedding table (9.2 → 4.3 GB; memory 16.1 → 14.8 GB). Same answers as v1 on our checks. Needs onw 0.3 or newer.
|
| 41 |
+
- v1: `hf download ryugyosoft/gemma-4-E4B-it-onw --revision v1 --local-dir ...`
|
| 42 |
+
|
| 43 |
+
Most of the download saving comes from storing identical weights once; in memory the NPU keeps one compiled copy per block size, so the footprint shrinks less.
|
| 44 |
+
|
| 45 |
+
## How it runs
|
| 46 |
+
|
| 47 |
+
- The static Gemma 4 decoder (head_dim-512 attention split into 256-wide heads, host-built masks) is imported as is.
|
| 48 |
+
- The token embedding and the 2.7 GB per-layer embedding table are looked up on the host (table reads: no NPU memory, no worker process).
|
| 49 |
+
- The LM head is one shared INT8 input for all block sizes; the logit softcap is arranged to avoid an NPU compiler bug.
|
| 50 |
+
- Details: [onw README](https://huggingface.co/ryugyosoft/onw).
|
| 51 |
+
|
| 52 |
+
## Files
|
| 53 |
+
|
| 54 |
+
| file | what |
|
| 55 |
+
|---|---|
|
| 56 |
+
| `seg00_S{1,16,64}.xml` + `seg00.bin` | static decoder for 1 / 16 / 64-token blocks (shared weights) |
|
| 57 |
+
| `seg01_S*.xml` + `seg01.bin` | LM head + logit softcapping |
|
| 58 |
+
| `vision.xml` | static vision encoder (280 soft tokens) |
|
| 59 |
+
| `shared.bin` | INT8 token embedding (= tied head), INT4 per-layer embedding table + id map |
|
| 60 |
+
| `engine.json`, tokenizer / config files | onw metadata, the original repo's configs |
|
| 61 |
+
|
| 62 |
+
## Quality
|
| 63 |
+
|
| 64 |
+
Channel-wise / group-wise INT4 (round-to-nearest): on our checks the first token matches the bf16 model; answers are fluent, word choices may diverge after a few tokens.
|
| 65 |
+
|
| 66 |
+
## License
|
| 67 |
+
|
| 68 |
+
Apache 2.0, same as the base model. The weights are re-quantized / restructured from it.
|