Download README_en.md from ryugyosoft/Qwen3.6-REAP-18B-A3B-onw: direct link, hf CLI and curl.
- Browser
- Download file 5.16 kB
-
https://huggingface.co/ryugyosoft/Qwen3.6-REAP-18B-A3B-onw/resolve/main/README_en.md
- Command line
-
hf download hf://ryugyosoft/Qwen3.6-REAP-18B-A3B-onw/README_en.md
-
curl -L -o README_en.md https://huggingface.co/ryugyosoft/Qwen3.6-REAP-18B-A3B-onw/resolve/main/README_en.md
Qwen3.6-REAP-18B-A3B for onw — entirely on the Intel NPU (text + image)
日本語 | English
Converted for onw (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.
Qwen/Qwen3.6-35B-A3B with half of its experts pruned (256 → 128 per layer, REAP). ~18B total / ~3B active parameters, text + image. The unpruned model: ryugyosoft/Qwen3.6-35B-A3B-onw.
To our knowledge, onw is the first public implementation that runs an MoE + Gated DeltaNet model, image input included, entirely on the NPU (survey by Fable 5.1, September 2026). This is the version pruned for 16 GB machines.
| NPU 3720 (Core Ultra 9 285HX) | |
|---|---|
| input | text + image |
| decode | 6.8 tok/s |
| prompt processing | image + question, 281 tokens: vision 0.5 s + prefill 15 s |
| download | 9.6 GB |
| memory in use (working set after loading) | ~12 GB |
| first start (NPU compile) / later | ~7 min / ~20 s |
Supported PCs: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory.
Use
Install onw (the engine) if you have not yet - one line, see Install in onw. On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the
curl ... | bashline in the onw README).irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iexIn the onw window (from the onw icon in the task tray), Models tab: pick Qwen3.6-REAP-18B-A3B and press Download (9.6 GB).
Server tab: pick it and press Load. The first load prepares it for the NPU (~7 min).
"Chat" to try it. Other apps can use it as an OpenAI-compatible API (base URL
http://localhost:8000/v1).

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
r = c.chat.completions.create(model="Qwen3.6-REAP-18B-A3B-onw", messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)
Advanced: hf download ryugyosoft/Qwen3.6-REAP-18B-A3B-onw and onw serve <folder> works too. Do not git clone without Git LFS (you would get pointer files instead of the weights; onw detects that and stops).
Pruning (REAP)
| this model (50% pruned) | unpruned | |
|---|---|---|
| download | 9.6 GB | 17 GB |
| memory in use (onw 0.7) | ~12 GB | ~20 GB |
| perplexity (chat data, bf16 reference) | 14.30 (+8.9%) | 13.13 |
| experts per layer | 128 (8 + shared active) | 256 (8 + shared active) |
REAP (Router-weighted Expert Activation Pruning): each expert's saliency is the mean, over the tokens routed to it, of gate weight x ||expert output||, measured on 16k tokens of chat-formatted instruction data in Japanese, English and code (databricks-dolly-15k-ja, databricks-dolly-15k, CodeAlpaca-20k). In every layer the 128 least salient experts are removed with their router rows; the softmax spreads over the survivors. No retraining. Held-out perplexity: 13.13 → 13.31 (75% kept) → 14.30 (50%).
Expect more loss on domains and languages the calibration mix does not cover. ~12 GB (onw 0.7) should fit a 16 GB PC (not measured on one yet). Reproduce with python -m onw.prune calibrate / eval and onw convert --prune saliency.json --keep 0.5.
How it runs
- Same structure as the unpruned model (41 segments + LM head, the host binds the chosen experts).
- With 128 experts, 16-token prompt blocks keep every expert bound once.
- Details: onw technical notes.
Files
| file | what |
|---|---|
seg*_S1.xml, seg*_S16.xml + seg*.bin |
41 segments + LM head (one weights file per segment) |
experts.bin (+ .json) |
40 layers x 128 experts, channel-wise INT4 |
vision.xml |
INT8 vision tower (512x512 -> 256 tokens) |
shared.bin |
INT8 LM head, INT4 token embedding |
engine.json, tokenizer / config files |
onw metadata, the original repo's configs |
Quality
Channel-wise / group-wise INT4 (round-to-nearest): on our checks the first token matches the bf16 model; answers are fluent, word choices may diverge after a few tokens.
License
Apache 2.0, same as the base model. The weights are pruned, re-quantized and restructured from it.