sakamakismile commited on
Commit
91154d9
·
verified ·
1 Parent(s): 5d7c613

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +209 -0
README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ base_model: ornith-ai/Ornith-1.5-397B
4
+ base_model_relation: quantized
5
+ quantized_by: Lna-Lab
6
+ library_name: gguf
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - gguf
10
+ - imatrix
11
+ - iq3_xxs
12
+ - moe
13
+ - qwen3_5_moe
14
+ - llama.cpp
15
+ - agentic-coding
16
+ language:
17
+ - en
18
+ - zh
19
+ - ja
20
+ ---
21
+
22
+ # Ornith-1.5-397B — IQ3_XXS GGUF
23
+
24
+ An **IQ3_XXS** quantization of [ornith-ai/Ornith-1.5-397B](https://huggingface.co/ornith-ai/Ornith-1.5-397B).
25
+
26
+ The official [Ornith-1.5-397B-GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-397B-GGUF) repository
27
+ stops at **Q4_K_M (224.08 GiB)**. That does not fit in 192 GB of VRAM. This one does.
28
+
29
+ | | size | BPW |
30
+ |---|---|---|
31
+ | official Q4_K_M | 224.08 GiB | 4.86 |
32
+ | **this IQ3_XXS** | **142.68 GiB** | **3.09** |
33
+
34
+ Split into 4 files of ≤45 GB. Point `llama.cpp` at `-00001-of-00004` and it loads all four.
35
+
36
+ ---
37
+
38
+ ## Files
39
+
40
+ | file | bytes |
41
+ |---|---|
42
+ | `Ornith-1.5-397B-IQ3_XXS-00001-of-00004.gguf` | 44,460,551,168 |
43
+ | `Ornith-1.5-397B-IQ3_XXS-00002-of-00004.gguf` | 44,665,664,288 |
44
+ | `Ornith-1.5-397B-IQ3_XXS-00003-of-00004.gguf` | 44,695,059,584 |
45
+ | `Ornith-1.5-397B-IQ3_XXS-00004-of-00004.gguf` | 19,373,124,448 |
46
+
47
+ Total 1098 tensors, 153,194,399,488 bytes across 4 files (142.67 GiB; the unsplit file is 153,194,399,008 B — the difference is per-split headers). See `SHA256SUMS.txt`.
48
+
49
+ For vision, use `mmproj-Ornith-1.5-397B-BF16.gguf` from the
50
+ [official GGUF repo](https://huggingface.co/ornith-ai/Ornith-1.5-397B-GGUF) (not mirrored here).
51
+
52
+ ---
53
+
54
+ ## How it was made
55
+
56
+ Source was the **official Q8_0 GGUF** (392.56 GiB, 8.51 BPW), not the BF16 checkpoint.
57
+ That means `--allow-requantize` was used — this is a re-quantization of an already-quantized
58
+ tensor set. Q8_0 is close to lossless, but this is stated plainly so you can weigh it.
59
+
60
+ ```
61
+ llama-quantize --allow-requantize \
62
+ --imatrix <imatrix.gguf> \
63
+ --token-embedding-type q5_K \
64
+ Ornith-1.5-397B-Q8_0.gguf Ornith-1.5-397B-IQ3_XXS.gguf IQ3_XXS 64
65
+
66
+ llama-gguf-split --split --split-max-size 45G \
67
+ Ornith-1.5-397B-IQ3_XXS.gguf Ornith-1.5-397B-IQ3_XXS
68
+ ```
69
+
70
+ `--token-embedding-type q5_K` overrides the IQ3_XXS default (`iq3_s`) for `token_embd`.
71
+ With a 248,320-token vocabulary carrying CJK, the extra ~250 MiB is worth it.
72
+
73
+ Quantization took **31m40s** on a Threadripper PRO 9985WX (64 cores, 64 threads).
74
+
75
+ ### About the importance matrix
76
+
77
+ **The imatrix is not ours and is not mirrored here.** We used
78
+ [`unsloth/Qwen3.5-397B-A17B-GGUF`](https://huggingface.co/unsloth/Qwen3.5-397B-A17B-GGUF)'s
79
+ `imatrix_unsloth.gguf_file` (80 chunks × 11264 tokens).
80
+
81
+ This works because **Ornith-1.5-397B is a light fine-tune of Qwen/Qwen3.5-397B-A17B**:
82
+
83
+ - the 1371 non-MTP tensor names are **identical sets** (set difference is empty)
84
+ - the vision tower is **bit-identical** (frozen), as are `linear_attn.A_log` and `dt_bias`
85
+ - the language trunk has cosine similarity **0.9993–0.99999** (relative L2 of 1–4%)
86
+ - the safetensors `total_size` differs by exactly 13,191,153,536 B — precisely the MTP head
87
+
88
+ We verified name compatibility before quantizing: **765 of 765 imatrix entries match tensors in
89
+ the Ornith Q8_0** (100%). The 180 quantizable tensors without imatrix coverage are norms and
90
+ `ssm_conv1d`, which are not quantized anyway.
91
+
92
+ Notably, 765 is the same `quantize.imatrix.entries_count` recorded in the official Ornith GGUF
93
+ headers — the official build used the same number of entries.
94
+
95
+ If you want a purpose-built imatrix, compute one against this model directly. We did not,
96
+ and we say so rather than implying otherwise.
97
+
98
+ ---
99
+
100
+ ## Measured
101
+
102
+ Pure CPU, Threadripper PRO 9985WX, 64 threads, `-dev none`:
103
+
104
+ | | prefill | decode |
105
+ |---|---|---|
106
+ | Q8_0 (reference) | 41.0–41.8 t/s | 9.7–9.8 t/s |
107
+ | **IQ3_XXS** | **33.9–34.6 t/s** | **13.0–13.1 t/s** |
108
+
109
+ Same prompt (three summer haiku, different kigo, one line each), `temp 0.8`, thinking off:
110
+
111
+ **Q8_0** —
112
+ ```
113
+ 金魚売り通り過ぎていく水の音
114
+ 青トマトかじれば夏の朝の味
115
+ 夕立やアスファルト跳ねる子らの声
116
+ ```
117
+
118
+ **IQ3_XXS** —
119
+ ```
120
+ 夏日や池の鯉ゆく水草かげ
121
+ 夏炉や炉の灰に眠る火の粉かな
122
+ 夏空や雲の切れ間より富士の山
123
+ ```
124
+
125
+ Both hold 5-7-5 and use three distinct summer kigo. Q8_0 reaches for more modern imagery,
126
+ IQ3_XXS sits closer to classical form. Neither is broken.
127
+
128
+ Perplexity has **not** been measured. Stated as missing rather than guessed at.
129
+
130
+ ### Why not IQ2
131
+
132
+ We also baked IQ2_XXS (97.65 GiB, 2.12 BPW) and **do not recommend it**. It answers factual
133
+ questions correctly ("日本の首都は東京です") but cannot carry out multi-step generation — asked
134
+ for haiku it emits bullet-point glossaries of season words, and at `temp 0.8` it degenerates into
135
+ repetition with stray tokens. At 2.12 BPW this model does not survive. It is not published here.
136
+
137
+ IQ3_XXS is, in our measurements, the floor.
138
+
139
+ ---
140
+
141
+ ## Usage
142
+
143
+ ```bash
144
+ llama-server -m Ornith-1.5-397B-IQ3_XXS-00001-of-00004.gguf \
145
+ -c 32768 --threads 64
146
+ ```
147
+
148
+ **Ornith is a reasoning model and it thinks at length.** With `-n 1500` it had not finished
149
+ deliberating. For direct answers:
150
+
151
+ ```
152
+ --chat-template-kwargs '{"enable_thinking":false}'
153
+ ```
154
+
155
+ If you keep thinking on, budget generously (the 35B sibling needed ≥6500 tokens) and strip
156
+ everything before `</think>` before parsing code out of a response — otherwise you will grade
157
+ the model's scratch work instead of its answer.
158
+
159
+ ### ⚠️ GPU offload does not work yet on SM 12.0
160
+
161
+ On 12× RTX PRO 2000 Blackwell (SM 12.0, CUDA 13.2) this model **crashes on GPU**:
162
+
163
+ ```
164
+ ggml_cuda_compute_forward: SOFT_MAX failed
165
+ CUDA error: invalid argument
166
+ ```
167
+
168
+ Isolated by bisecting `-ngl`:
169
+
170
+ - `-ngl 1` (layer 59, a **full_attention** layer) → runs
171
+ - `-ngl 2` (adds layer 58, a **linear_attention** layer) → crashes
172
+
173
+ So it is the linear-attention (gated delta net) path. `-fa on` does not help
174
+ (`flash_attn = enabled` is logged and SOFT_MAX is still reached), nor does `--no-warmup`,
175
+ nor `-ub 1 -b 1`. Reproduced on both a 2026-08-10 build and on master at `d59d455`
176
+ (174 commits newer). CPU inference is unaffected.
177
+
178
+ Separately, llama.cpp **misclassifies Blackwell as an integrated GPU** because
179
+ `cudaDeviceProp.integrated` is non-zero (the driver API correctly reports 0 for the same device).
180
+ Only the first "iGPU" is kept, so `-sm`/`-ts` silently do nothing and everything piles onto
181
+ device 0. Upstream [#26901](https://github.com/ggml-org/llama.cpp/issues/26901), open since
182
+ 2026-08-11. Work around it by naming devices explicitly:
183
+
184
+ ```
185
+ -dev CUDA0,CUDA1,CUDA2,CUDA3,CUDA4,CUDA5,CUDA6,CUDA7,CUDA8,CUDA9,CUDA10,CUDA11
186
+ -ts 4.5,5,5,5,5,5,5,5,5,5,5,6.5
187
+ ```
188
+
189
+ That does distribute the layers correctly (verified in the load log) — the SOFT_MAX crash is a
190
+ separate, unresolved problem.
191
+
192
+ ### Note for anyone re-converting from safetensors
193
+
194
+ `config.json` declares `mtp_num_hidden_layers=1`, but **there is not a single MTP tensor in the
195
+ checkpoint** (1371 tensors, 0 MTP) or in the official GGUF (1098 tensors, 0 nextn). The 35B-A3B
196
+ sibling does ship 785 of them; the 397B does not, in either 1.0 or 1.5.
197
+
198
+ Convert with `--no-mtp`. Without it you get a GGUF declaring `block_count=61` with an empty
199
+ `blk.60`, and llama.cpp fails at load with a missing-tensor error.
200
+
201
+ ---
202
+
203
+ ## Attribution
204
+
205
+ - Base model: [ornith-ai/Ornith-1.5-397B](https://huggingface.co/ornith-ai/Ornith-1.5-397B) — MIT. All credit for the model belongs to its authors.
206
+ - Importance matrix: [unsloth/Qwen3.5-397B-A17B-GGUF](https://huggingface.co/unsloth/Qwen3.5-397B-A17B-GGUF).
207
+ - Tooling: [llama.cpp](https://github.com/ggml-org/llama.cpp).
208
+
209
+ This repository contributes quantized weights and the measurements above. Nothing else.