# granite-4.2-3b for onw — entirely on the Intel NPU
[日本語](README.md) | **English**
[ibm-granite/granite-4.2-3b](https://huggingface.co/ibm-granite/granite-4.2-3b) converted for **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.
IBM's [Granite 4.2](https://huggingface.co/collections/ibm-granite/granite-42-language-models) adds native thinking and reasoning-augmented tool calling, trained with agentic RL (code fixes, terminal use, web search; trajectories of up to 200 tool calls). A plain (Llama-type) transformer. Japanese is officially supported but its multilingual scores are lower than English: Japanese answers can contain wrong kanji or odd content (for Japanese, [llm-jp-4.1-8b-thinking](https://huggingface.co/ryugyosoft/llm-jp-4.1-8b-thinking-onw) is the better pick). This 3B is the lightest Granite 4.2 and among the fastest onw models (its card: BFCL v4 52.4, the same as the 8B; τ³-bench 51.0; IFBench 74.3).
**Thinking and tools**: onw answers without thinking by default (the chat page's thinking switch, or `chat_template_kwargs: {"enable_thinking": true}` in the API; the thinking comes back as `reasoning_content`). Tools via OpenAI-style `tools` (parallel calls too). Thinking runs long: use `max_tokens` of 2000+ with tools. The original card recommends temperature 1.0, top_p 0.95 (onw defaults to greedy).
| NPU 3720 (Core Ultra 9 285HX) | |
|---|---|
| input | text |
| decode | **6.7 tok/s (~7.7 in chat with prompt lookup)** |
| prompt processing | under 1 s for a short question |
| download | **2.1 GB** |
| memory in use (working set after loading) | ~6 GB (6.4 GB measured on a 16 GB Lunar Lake; with the memory-saving mode in the settings ~2.8 GB idle and while answering, 5.1 GB for the ~3 s of a swap) |
| first start (NPU compile) / later | ~6 min / ~20 s |
**Supported PCs**: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory.
## Use
1. **Install onw (the engine)** if you have not yet - one line, see [Install in onw](https://huggingface.co/ryugyosoft/onw). On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the `curl ... | bash` line in the onw README).
```powershell
irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iex
```
2. In the **onw window** (from the onw icon in the task tray), Models tab: pick **granite-4.2-3b** and press Download (2.1 GB).
3. Server tab: pick it and press Load. The first load prepares it for the NPU (~6 min).
4. "Chat" to try it. Other apps can use it as an **OpenAI-compatible API** (base URL `http://localhost:8000/v1`).
```python
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
r = c.chat.completions.create(model="granite-4.2-3b-onw", messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)
```
Advanced: `hf download ryugyosoft/granite-4.2-3b-onw` and `onw serve ` works too. Do not `git clone` without Git LFS (you would get pointer files instead of the weights; onw detects that and stops).
## How it runs
- 40 layers in 5 segments of 8 plus an LM-head segment; Granite's embedding / residual / attention / logits multipliers are built into the graphs.
- Plain group-128 INT4 loses noticeable accuracy on Granite, so the weights use a **sign-aware INT4** (the scale fits max/7 or -min/8, same layout and speed as plain INT4), with only the attention q/k/v in group-128 INT8 (a small speed cost).
- The LM head is its own INT8 segment; rows whose scale would be an fp16 subnormal (zero on the NPU) are re-encoded.
- Details: [onw technical notes](https://huggingface.co/ryugyosoft/onw/blob/main/TECHNICAL_en.md).
## Files
| file | what |
|---|---|
| `seg*_S1.xml`, `seg*_S16.xml` + `seg*.bin` | 5 segments + LM head (one weights file per segment) |
| `shared.bin` | INT8 LM head, INT4 token embedding (host lookup) |
| `engine.json`, tokenizer / config files | onw metadata, the original repo's configs |
## Quality
Teacher-forcing a 400-token bf16 answer: 344/400 tokens match (316 with plain INT4). The NPU and the CPU (fp32) agree; first tokens of 20 Japanese prompts: 15/20. Checked: two rounds of tool calls through the final answer, the thinking split.
## License
Apache 2.0, same as the base model. The weights are re-quantized / restructured from it; no training.