# NeoHorse-1-4B for onw — entirely on the Intel NPU [日本語](README.md) | **English** [TokenRhythm/NeoHorse-1-4B](https://huggingface.co/TokenRhythm/NeoHorse-1-4B) converted for **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download. TokenRhythm's [NeoHorse-1-4B](https://huggingface.co/TokenRhythm/NeoHorse-1-4B) is [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) trained further for agent harnesses, tool calls, coding and instruction following (its card: +6 to +10 points over Qwen3.5-4B on agent benchmarks, ten-benchmark average 58.9 -> 64.9). Same Gated DeltaNet + gated attention structure as Qwen3.5, 4B, text only. A 2.4 GB download and ~6 GB of memory: plenty of headroom on 16 GB PCs. **Tool calls**: OpenAI-style `tools` (parallel calls too). In thinking mode the thinking comes back as `reasoning_content`. The original card recommends temperature 1.0, top_p 0.95, top_k 20 (onw defaults to greedy). | NPU 3720 (Core Ultra 9 285HX) | | |---|---| | input | text | | decode | **6.3-6.5 tok/s (~1.6x Qwen3.5-9B on the same NPU)** | | prompt processing | 23 tokens in ~0.8 s | | download | **2.4 GB** | | memory in use (working set after loading) | ~6 GB | | first start (NPU compile) / later | ~6 min / ~20 s | **Supported PCs**: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory. ## Use 1. **Install onw (the engine)** if you have not yet - one line, see [Install in onw](https://huggingface.co/ryugyosoft/onw). On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the `curl ... | bash` line in the onw README). ```powershell irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iex ``` 2. In the **onw window** (from the onw icon in the task tray), Models tab: pick **NeoHorse-1-4B** and press Download (2.4 GB). 3. Server tab: pick it and press Load. The first load prepares it for the NPU (~6 min). 4. "Chat" to try it. Other apps can use it as an **OpenAI-compatible API** (base URL `http://localhost:8000/v1`). The Models tab of the onw window Server tab Settings tab ```python from openai import OpenAI c = OpenAI(base_url="http://localhost:8000/v1", api_key="none") r = c.chat.completions.create(model="NeoHorse-1-4B-onw", messages=[{"role": "user", "content": "Hello"}]) print(r.choices[0].message.content) ``` Advanced: `hf download ryugyosoft/NeoHorse-1-4B-onw` and `onw serve ` works too. Do not `git clone` without Git LFS (you would get pointer files instead of the weights; onw detects that and stops). ## How it runs - 32 layers (24 Gated DeltaNet + 8 gated attention) in 4 segments of 8 layers, group-128 INT4; the LM head shares the token embedding (as in the original) and is skipped for prompt blocks that need no logits. - DeltaNet: 1-token matrix form for decoding; 16-token prompt blocks in chunkwise-parallel form with an exact block-doubling inverse. The DeltaNet output is scaled by 1024 (fp16 subnormals) and the gated RMSNorm prescaled so it cannot overflow fp16. - Details: [onw technical notes](https://huggingface.co/ryugyosoft/onw/blob/main/TECHNICAL_en.md). ## Files | file | what | |---|---| | `seg*_S1.xml`, `seg*_S16.xml` + `seg*.bin` | 4 segments + LM head (one weights file per segment) | | `shared.bin` | INT8 LM head (tied to the embedding), INT4 token embedding (host lookup) | | `engine.json`, tokenizer / config files | onw metadata, the original repo's configs | ## Quality Group-128 INT4 segments, INT8 LM head, INT4 token embedding. Teacher-forced against the bf16 model on the CPU: 24/24 tokens with thinking, and everything up to the end of the answer (`<|im_end|>`) without. First tokens of 20 Japanese prompts: 18/20 match (the other two are ties or near-ties in the bf16 model too), identical on the NPU and the CPU (fp32). Checked: two parallel tool calls and an answer built from their results, the thinking split. ## License Apache 2.0, same as the base model (its LICENSE, with Qwen's copyright notice, is included). The weights are re-quantized / restructured from it; no training.