Qwen3.6-REAP-18B-A3B-onw / README_en.md
ryugyosoft's picture
Card: supported PCs incl. Panther Lake
e6a9b3b verified
|
Raw History Blame Contribute Delete
5.16 kB

Qwen3.6-REAP-18B-A3B for onw — entirely on the Intel NPU (text + image)

日本語 | English

Converted for onw (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.

Qwen/Qwen3.6-35B-A3B with half of its experts pruned (256 → 128 per layer, REAP). ~18B total / ~3B active parameters, text + image. The unpruned model: ryugyosoft/Qwen3.6-35B-A3B-onw.

To our knowledge, onw is the first public implementation that runs an MoE + Gated DeltaNet model, image input included, entirely on the NPU (survey by Fable 5.1, September 2026). This is the version pruned for 16 GB machines.

NPU 3720 (Core Ultra 9 285HX)
input text + image
decode 6.8 tok/s
prompt processing image + question, 281 tokens: vision 0.5 s + prefill 15 s
download 9.6 GB
memory in use (working set after loading) ~12 GB
first start (NPU compile) / later ~7 min / ~20 s

Supported PCs: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory.

Use

  1. Install onw (the engine) if you have not yet - one line, see Install in onw. On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the curl ... | bash line in the onw README).

    irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iex
    
  2. In the onw window (from the onw icon in the task tray), Models tab: pick Qwen3.6-REAP-18B-A3B and press Download (9.6 GB).

  3. Server tab: pick it and press Load. The first load prepares it for the NPU (~7 min).

  4. "Chat" to try it. Other apps can use it as an OpenAI-compatible API (base URL http://localhost:8000/v1).

The Models tab of the onw window

Server tab Settings tab

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
r = c.chat.completions.create(model="Qwen3.6-REAP-18B-A3B-onw", messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)

Advanced: hf download ryugyosoft/Qwen3.6-REAP-18B-A3B-onw and onw serve <folder> works too. Do not git clone without Git LFS (you would get pointer files instead of the weights; onw detects that and stops).

Pruning (REAP)

this model (50% pruned) unpruned
download 9.6 GB 17 GB
memory in use (onw 0.7) ~12 GB ~20 GB
perplexity (chat data, bf16 reference) 14.30 (+8.9%) 13.13
experts per layer 128 (8 + shared active) 256 (8 + shared active)

REAP (Router-weighted Expert Activation Pruning): each expert's saliency is the mean, over the tokens routed to it, of gate weight x ||expert output||, measured on 16k tokens of chat-formatted instruction data in Japanese, English and code (databricks-dolly-15k-ja, databricks-dolly-15k, CodeAlpaca-20k). In every layer the 128 least salient experts are removed with their router rows; the softmax spreads over the survivors. No retraining. Held-out perplexity: 13.13 → 13.31 (75% kept) → 14.30 (50%).

Expect more loss on domains and languages the calibration mix does not cover. ~12 GB (onw 0.7) should fit a 16 GB PC (not measured on one yet). Reproduce with python -m onw.prune calibrate / eval and onw convert --prune saliency.json --keep 0.5.

How it runs

  • Same structure as the unpruned model (41 segments + LM head, the host binds the chosen experts).
  • With 128 experts, 16-token prompt blocks keep every expert bound once.
  • Details: onw technical notes.

Files

file what
seg*_S1.xml, seg*_S16.xml + seg*.bin 41 segments + LM head (one weights file per segment)
experts.bin (+ .json) 40 layers x 128 experts, channel-wise INT4
vision.xml INT8 vision tower (512x512 -> 256 tokens)
shared.bin INT8 LM head, INT4 token embedding
engine.json, tokenizer / config files onw metadata, the original repo's configs

Quality

Channel-wise / group-wise INT4 (round-to-nearest): on our checks the first token matches the bf16 model; answers are fluent, word choices may diverge after a few tokens.

License

Apache 2.0, same as the base model. The weights are pruned, re-quantized and restructured from it.