NeoHorse-1-4B-onw / README_en.md
ryugyosoft's picture
Add files using upload-large-folder tool
1f5d799 verified
|
Raw History Blame Contribute Delete
4.74 kB

NeoHorse-1-4B for onw — entirely on the Intel NPU

日本語 | English

TokenRhythm/NeoHorse-1-4B converted for onw (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download.

TokenRhythm's NeoHorse-1-4B is Qwen/Qwen3.5-4B trained further for agent harnesses, tool calls, coding and instruction following (its card: +6 to +10 points over Qwen3.5-4B on agent benchmarks, ten-benchmark average 58.9 -> 64.9). Same Gated DeltaNet + gated attention structure as Qwen3.5, 4B, text only. A 2.4 GB download and ~6 GB of memory: plenty of headroom on 16 GB PCs.

Tool calls: OpenAI-style tools (parallel calls too). In thinking mode the thinking comes back as reasoning_content. The original card recommends temperature 1.0, top_p 0.95, top_k 20 (onw defaults to greedy).

NPU 3720 (Core Ultra 9 285HX)
input text
decode 6.3-6.5 tok/s (~1.6x Qwen3.5-9B on the same NPU)
prompt processing 23 tokens in ~0.8 s
download 2.4 GB
memory in use (working set after loading) ~6 GB
first start (NPU compile) / later ~6 min / ~20 s

Supported PCs: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory.

Use

  1. Install onw (the engine) if you have not yet - one line, see Install in onw. On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the curl ... | bash line in the onw README).

    irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iex
    
  2. In the onw window (from the onw icon in the task tray), Models tab: pick NeoHorse-1-4B and press Download (2.4 GB).

  3. Server tab: pick it and press Load. The first load prepares it for the NPU (~6 min).

  4. "Chat" to try it. Other apps can use it as an OpenAI-compatible API (base URL http://localhost:8000/v1).

The Models tab of the onw window

Server tab Settings tab

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
r = c.chat.completions.create(model="NeoHorse-1-4B-onw", messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)

Advanced: hf download ryugyosoft/NeoHorse-1-4B-onw and onw serve <folder> works too. Do not git clone without Git LFS (you would get pointer files instead of the weights; onw detects that and stops).

How it runs

  • 32 layers (24 Gated DeltaNet + 8 gated attention) in 4 segments of 8 layers, group-128 INT4; the LM head shares the token embedding (as in the original) and is skipped for prompt blocks that need no logits.
  • DeltaNet: 1-token matrix form for decoding; 16-token prompt blocks in chunkwise-parallel form with an exact block-doubling inverse. The DeltaNet output is scaled by 1024 (fp16 subnormals) and the gated RMSNorm prescaled so it cannot overflow fp16.
  • Details: onw technical notes.

Files

file what
seg*_S1.xml, seg*_S16.xml + seg*.bin 4 segments + LM head (one weights file per segment)
shared.bin INT8 LM head (tied to the embedding), INT4 token embedding (host lookup)
engine.json, tokenizer / config files onw metadata, the original repo's configs

Quality

Group-128 INT4 segments, INT8 LM head, INT4 token embedding. Teacher-forced against the bf16 model on the CPU: 24/24 tokens with thinking, and everything up to the end of the answer (<|im_end|>) without. First tokens of 20 Japanese prompts: 18/20 match (the other two are ties or near-ties in the bf16 model too), identical on the NPU and the CPU (fp32). Checked: two parallel tool calls and an answer built from their results, the thinking split.

License

Apache 2.0, same as the base model (its LICENSE, with Qwen's copyright notice, is included). The weights are re-quantized / restructured from it; no training.