|
Download README_en.md from ryugyosoft/Ornith-1.5-9B-onw: direct link, hf CLI and curl.
- Browser
- Download file 4.99 kB
-
https://huggingface.co/ryugyosoft/Ornith-1.5-9B-onw/resolve/main/README_en.md
- Command line
-
hf download hf://ryugyosoft/Ornith-1.5-9B-onw/README_en.md
-
curl -L -o README_en.md https://huggingface.co/ryugyosoft/Ornith-1.5-9B-onw/resolve/main/README_en.md
4.99 kB
| # Ornith-1.5-9B for onw — entirely on the Intel NPU (text + image) | |
| [日本語](README.md) | **English** | |
| [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) converted for **[onw](https://huggingface.co/ryugyosoft/onw)** (俺のNPUがこんなに動くわけない - "there's no way my NPU runs this well"), the engine that runs LLMs entirely on the Intel NPU. This repo holds only the model; the engine is a separate download. | |
| ornith-ai's [Ornith-1.5](https://huggingface.co/ornith-ai) takes Qwen3.5-family models through continued pre-training, mid-training and post-training with a self-improvement loop aimed at coding agents, tool calls and long tasks (its card reports large gains over the base Qwen on SWE-bench Verified, Terminal-Bench, MCP-Atlas and more). The architecture is Qwen3.5 / 3.6's, so onw's Qwen conversion and runtime apply as they are. This 9B is dense, text + image. | |
| **Tool calls**: OpenAI-style `tools` (parallel calls too). In thinking mode the thinking comes back as `reasoning_content`; send it back in the history and the model gets it (Ornith's template keeps earlier turns' thinking, which agents rely on). The original card recommends temperature 1.0, top_p 0.95, top_k 20, presence_penalty 1.5 for general use and temperature 0.6 for coding (onw defaults to greedy). | |
| | NPU 3720 (Core Ultra 9 285HX) | | | |
| |---|---| | |
| | input | text + image | | |
| | decode | **4.9 tok/s (9-15 tok/s with prompt lookup)** | | |
| | prompt processing | image + question, 281 tokens: vision 0.5 s + prefill 8.0 s | | |
| | download | **5.3 GB** | | |
| | memory in use (working set after loading) | ~11 GB | | |
| | first start (NPU compile) / later | ~10 min / ~20 s | | |
| **Supported PCs**: Intel Core Ultra with an NPU (series 1 / 2: Meteor Lake, Arrow Lake, Lunar Lake; series 3: Panther Lake), Windows 11 or Ubuntu 22.04+. The numbers above are measured on the test machine (NPU 3720) and vary with the NPU and memory. | |
| ## Use | |
| 1. **Install onw (the engine)** if you have not yet - one line, see [Install in onw](https://huggingface.co/ryugyosoft/onw). On Windows open "PowerShell" from the Start menu, paste this and press Enter (Ubuntu: the `curl ... | bash` line in the onw README). | |
| ```powershell | |
| irm https://huggingface.co/ryugyosoft/onw/resolve/main/install.ps1 | iex | |
| ``` | |
| 2. In the **onw window** (from the onw icon in the task tray), Models tab: pick **Ornith-1.5-9B** and press Download (5.3 GB). | |
| 3. Server tab: pick it and press Load. The first load prepares it for the NPU (~10 min). | |
| 4. "Chat" to try it. Other apps can use it as an **OpenAI-compatible API** (base URL `http://localhost:8000/v1`). | |
| <img src="https://huggingface.co/ryugyosoft/onw/resolve/main/images/onw-models.png" width="640" alt="The Models tab of the onw window"> | |
| <img src="https://huggingface.co/ryugyosoft/onw/resolve/main/images/onw-server.png" width="49%" alt="Server tab"> <img src="https://huggingface.co/ryugyosoft/onw/resolve/main/images/onw-settings.png" width="49%" alt="Settings tab"> | |
| ```python | |
| from openai import OpenAI | |
| c = OpenAI(base_url="http://localhost:8000/v1", api_key="none") | |
| r = c.chat.completions.create(model="Ornith-1.5-9B-onw", messages=[{"role": "user", "content": "Hello"}]) | |
| print(r.choices[0].message.content) | |
| ``` | |
| Advanced: `hf download ryugyosoft/Ornith-1.5-9B-onw` and `onw serve <folder>` works too. Do not `git clone` without Git LFS (you would get pointer files instead of the weights; onw detects that and stops). | |
| ## How it runs | |
| - 32 layers (24 Gated DeltaNet + 8 gated attention) in 4 segments of 8 layers, group-128 INT4, plus an LM-head segment skipped for prompt blocks that need no logits. | |
| - DeltaNet: 1-token matrix form for decoding; 16-token prompt blocks in chunkwise-parallel form with an exact block-doubling inverse. The DeltaNet output is scaled by 1024 (fp16 subnormals) and the gated RMSNorm prescaled so it cannot overflow fp16. | |
| - Vision: the 27-layer ViT as one static graph for 512x512 input; MRoPE positions for image tokens are computed on the host. | |
| - Details: [onw technical notes](https://huggingface.co/ryugyosoft/onw/blob/main/TECHNICAL_en.md). | |
| ## Files | |
| | file | what | | |
| |---|---| | |
| | `seg*_S1.xml`, `seg*_S16.xml` + `seg*.bin` | 4 segments + LM head (one weights file per segment) | | |
| | `vision.xml` | INT8 vision tower (512x512 -> 256 tokens) | | |
| | `shared.bin` | INT8 LM head, INT4 token embedding (host lookup) | | |
| | `engine.json`, tokenizer / config files | onw metadata, the original repo's configs | | |
| ## Quality | |
| Group-128 INT4 segments, INT8 LM head, INT4 token embedding (host lookup), INT8 vision tower. Teacher-forced against the bf16 model on the CPU, all 10 tokens up to the end of the answer (`<|im_end|>`) match. Checked: an image (shapes and colours), two parallel tool calls and an answer built from their results, the thinking split. | |
| ## License | |
| MIT, same as the base model. The weights are re-quantized / restructured from it; no training. | |