Download README_en.md from ryugyosoft/NeoHorse-1-4B-onw: direct link, hf CLI and curl.
- Browser
- Download file 4.74 kB
-
https://huggingface.co/ryugyosoft/NeoHorse-1-4B-onw/resolve/main/README_en.md
- Command line
-
hf download hf://ryugyosoft/NeoHorse-1-4B-onw/README_en.md
-
curl -L -o README_en.md https://huggingface.co/ryugyosoft/NeoHorse-1-4B-onw/resolve/main/README_en.md
NeoHorse-1-4B for onw — entirely on the Intel NPU
日本語 | English
TokenRhythm/NeoHorse-1-4B converted for onw (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.
TokenRhythm's NeoHorse-1-4B is Qwen/Qwen3.5-4B trained further for agent harnesses, tool calls, coding and instruction following (its card: +6 to +10 points over Qwen3.5-4B on agent benchmarks, ten-benchmark average 58.9 -> 64.9). Same Gated DeltaNet + gated attention structure as Qwen3.5, 4B, text only. A 2.4 GB download and ~6 GB of memory: plenty of headroom on 16 GB PCs.
Tool calls: OpenAI-style tools (parallel calls too). In thinking mode the thinking comes back as reasoning_content. The original card recommends temperature 1.0, top_p 0.95, top_k 20 (onw defaults to greedy).
| NPU 3720 (Core Ultra 9 285HX) | |
|---|---|
| input | text |
| decode | 6.3-6.5 tok/s (~1.6x Qwen3.5-9B on the same NPU) |
| prompt processing | 23 tokens in ~0.8 s |
| download | 2.4 GB |
| memory in use (working set after loading) | ~6 GB |
| first start (NPU compile) / later | ~6 min / ~20 s |
Supported PCs: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory.
Use
Install onw (the engine) if you have not yet - one line, see Install in onw. On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the
curl ... | bashline in the onw README).irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iexIn the onw window (from the onw icon in the task tray), Models tab: pick NeoHorse-1-4B and press Download (2.4 GB).
Server tab: pick it and press Load. The first load prepares it for the NPU (~6 min).
"Chat" to try it. Other apps can use it as an OpenAI-compatible API (base URL
http://localhost:8000/v1).

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
r = c.chat.completions.create(model="NeoHorse-1-4B-onw", messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)
Advanced: hf download ryugyosoft/NeoHorse-1-4B-onw and onw serve <folder> works too. Do not git clone without Git LFS (you would get pointer files instead of the weights; onw detects that and stops).
How it runs
- 32 layers (24 Gated DeltaNet + 8 gated attention) in 4 segments of 8 layers, group-128 INT4; the LM head shares the token embedding (as in the original) and is skipped for prompt blocks that need no logits.
- DeltaNet: 1-token matrix form for decoding; 16-token prompt blocks in chunkwise-parallel form with an exact block-doubling inverse. The DeltaNet output is scaled by 1024 (fp16 subnormals) and the gated RMSNorm prescaled so it cannot overflow fp16.
- Details: onw technical notes.
Files
| file | what |
|---|---|
seg*_S1.xml, seg*_S16.xml + seg*.bin |
4 segments + LM head (one weights file per segment) |
shared.bin |
INT8 LM head (tied to the embedding), INT4 token embedding (host lookup) |
engine.json, tokenizer / config files |
onw metadata, the original repo's configs |
Quality
Group-128 INT4 segments, INT8 LM head, INT4 token embedding. Teacher-forced against the bf16 model on the CPU: 24/24 tokens with thinking, and everything up to the end of the answer (<|im_end|>) without. First tokens of 20 Japanese prompts: 18/20 match (the other two are ties or near-ties in the bf16 model too), identical on the NPU and the CPU (fp32). Checked: two parallel tool calls and an answer built from their results, the thinking split.
License
Apache 2.0, same as the base model (its LICENSE, with Qwen's copyright notice, is included). The weights are re-quantized / restructured from it; no training.