Sebastian Aldrin commited on
Commit
c863d2a
Β·
verified Β·
1 Parent(s): 9f93a09

readme: add architecture + quant tables, keep voice

Browse files
Files changed (1) hide show
  1. README.md +52 -25
README.md CHANGED
@@ -23,32 +23,41 @@ language:
23
 
24
  # Nex-N2-mini IQ3_XXS
25
 
26
- imatrix-calibrated IQ3_XXS of nex-n2-mini. ~13gb, fits in 15gb GPU memory with room left for context. mostly made it because i wanted to run this model on a laptop iGPU and the existing quants were all 16gb+.
27
 
28
- gets ~14 tok/s on CPU only on a ryzen 7 pro 7735u, more with vulkan offload.
29
 
30
- ## you'll hit this if you quantize it yourself
31
 
32
- ```
33
- missing tensor 'blk.39.nextn.eh_proj.weight'
34
- ```
35
 
36
- took me forever to figure out. nex agi didnt release the MTP weights with the model, but the config says they're there, so the convert script writes "has MTP" into the GGUF header and then nothing loads it.
 
 
 
 
 
 
 
 
 
 
 
37
 
38
- two settings to patch:
39
 
40
- ```
41
- qwen35moe.nextn_predict_layers: 1 β†’ 0
42
- qwen35moe.block_count: 41 β†’ 40
43
- ```
44
-
45
- `patch_gguf.py` in the repo does it. 4-byte edits, doesnt shift anything else, takes 30s. way faster than re-converting from safetensors (8h).
46
 
47
  ## using it
48
 
49
- needs a llama.cpp from after 2026-02-10 (when qwen35moe arch landed in PR #19468).
50
 
51
- LM Studio: drop the gguf in `~/.lmstudio/models/<you>/Nex-N2-mini-GGUF/`, load it. update the bundled llama.cpp runtime to 2.13+ if loading fails.
52
 
53
  Ollama 0.19+:
54
  ```
@@ -65,27 +74,45 @@ llama-cli -m Nex-N2-mini-IQ3_XXS.gguf -ngl 999 \
65
 
66
  ollama before 0.19 wont work β€” too old to know the qwen35moe arch.
67
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68
  ## stuff to know
69
 
70
- - its a reasoning model so outputs have `<think>...</think>` blocks. handle that or strip it
71
- - no MTP speedup since the weights arent in the public release
72
- - text only. the base model's config.json has vision/video token slots but llama.cpp's qwen35moe converter is text-only (PR #19468 title literally says "no vision"). dont expect image support
73
- - Q3 means maybe 1-3% benchmark drop vs Q4. for chat/tool use i cant tell
74
 
75
  ## how i made it
76
 
77
  ```
78
  huggingface-cli download nex-agi/Nex-N2-mini --local-dir source
79
  python convert_hf_to_gguf.py source --outtype f16
80
- python patch_gguf.py source/*-F16.gguf # whatever the convert script named it
81
  llama-imatrix -m <f16.gguf> -f calibration.txt -o imatrix.dat --chunks 50
82
- llama-quantize --imatrix imatrix.dat <f16.gguf> out.gguf IQ3_XXS
83
  ```
84
 
85
- calibration text was Pride and Prejudice from gutenberg plus the nex-n2 README. imatrix-aware quantizer kept attention tensors at Q4_K and pushed expert FFN weights down to IQ3, ended up at 3.14 bpw avg.
86
 
87
- `imatrix.dat` is in the repo if you wanna re-quantize to something else.
88
 
89
  ---
90
 
91
- credit to nex agi for the model, Qwen team for the base arch, ggml-org for llama.cpp. apache 2.0, same as the base.
 
23
 
24
  # Nex-N2-mini IQ3_XXS
25
 
26
+ imatrix-calibrated IQ3_XXS of [nex-agi/Nex-N2-mini](https://huggingface.co/nex-agi/Nex-N2-mini). 13 GB, fits in 15 GB GPU memory with room for context. smallest quant of this model on the hub as of 2026-06-09.
27
 
28
+ made it because i wanted to run nex-n2-mini on my laptop's AMD iGPU (15 GB GTT cap) and every existing quant was 14 GB+.
29
 
30
+ gets **~14 tok/s on CPU only** (Ryzen 7 PRO 7735U, no GPU offload). vulkan offload pushes it higher.
31
 
32
+ ## architecture
 
 
33
 
34
+ | | |
35
+ |---|---|
36
+ | Base | Qwen3.5-35B-A3B-Base (post-trained by Nex AGI) |
37
+ | Architecture | `qwen35moe` |
38
+ | Total params | ~35B |
39
+ | Active params | ~3B per token |
40
+ | Experts | 256 total, 8 routed + 1 shared per token |
41
+ | Hidden size | 2048 |
42
+ | Trunk layers | 40 (MTP head not included β€” see below) |
43
+ | Train context | 262144 |
44
+ | Vocab | 248320 |
45
+ | Vision | not in this GGUF (text-only β€” see "stuff to know") |
46
 
47
+ ## file
48
 
49
+ | file | quant | size | bpw | notes |
50
+ |---|---|---:|---:|---|
51
+ | `Nex-N2-mini-IQ3_XXS.gguf` | IQ3_XXS | 12.7 GB | 3.14 | attention kept at Q4_K, FFN experts pushed to IQ3 |
52
+ | `imatrix.dat` | β€” | 183 MB | β€” | importance matrix, re-quantize from this if you want a different size |
53
+ | `patch_gguf.py` | β€” | 3.5 KB | β€” | fixes the MTP load error (see below) |
54
+ | `Modelfile` | β€” | 1 KB | β€” | for `ollama create` |
55
 
56
  ## using it
57
 
58
+ needs a llama.cpp from after **2026-02-10** (when qwen35moe arch landed in [PR #19468](https://github.com/ggml-org/llama.cpp/pull/19468)).
59
 
60
+ LM Studio: drop the gguf in `~/.lmstudio/models/<you>/Nex-N2-mini-GGUF/`, load it. update the bundled llama.cpp runtime to 2.13+ if it refuses to load.
61
 
62
  Ollama 0.19+:
63
  ```
 
74
 
75
  ollama before 0.19 wont work β€” too old to know the qwen35moe arch.
76
 
77
+ ## if you quantize it yourself
78
+
79
+ you'll hit:
80
+ ```
81
+ missing tensor 'blk.39.nextn.eh_proj.weight'
82
+ ```
83
+
84
+ took me forever to figure out. nex agi didn't release the MTP draft head weights with the public release, but `config.json` claims they exist, so the convert script writes "has MTP" into the GGUF header and llama.cpp's loader then refuses because it can't find the tensor.
85
+
86
+ two metadata values to flip:
87
+
88
+ ```
89
+ qwen35moe.nextn_predict_layers: 1 β†’ 0
90
+ qwen35moe.block_count: 41 β†’ 40
91
+ ```
92
+
93
+ `patch_gguf.py` in this repo does it. 4-byte edits, idempotent, takes 30 seconds. way faster than re-converting from safetensors (8h on a laptop).
94
+
95
  ## stuff to know
96
 
97
+ - **reasoning model** β€” outputs contain `<think>...</think>` blocks. handle them in your wrapper or strip them
98
+ - **no MTP speedup** β€” weights aren't in the public release. inference works fine, you just don't get the speculative-decoding bonus
99
+ - **text only** β€” the base model's `config.json` has `vision_config` + image/video token slots, but llama.cpp's qwen35moe converter is text-only (PR #19468 literally titled "no vision"). if you want vision, look at quants that ship an `mmproj-*.gguf` alongside
100
+ - **Q3 means ~1-3% benchmark drop vs Q4** β€” for chat and tool-calling i can't tell the difference. for code/math it'll be more noticeable
101
 
102
  ## how i made it
103
 
104
  ```
105
  huggingface-cli download nex-agi/Nex-N2-mini --local-dir source
106
  python convert_hf_to_gguf.py source --outtype f16
107
+ python patch_gguf.py source/*-F16.gguf
108
  llama-imatrix -m <f16.gguf> -f calibration.txt -o imatrix.dat --chunks 50
109
+ llama-quantize --imatrix imatrix.dat <f16.gguf> Nex-N2-mini-IQ3_XXS.gguf IQ3_XXS
110
  ```
111
 
112
+ calibration text = Pride and Prejudice from gutenberg + the nex-n2 README (~750 KB total, 50 chunks of 512 tokens). imatrix-aware quantizer kept attention tensors at Q4_K and pushed expert FFN weights down to IQ3 β€” ended up at 3.14 bpw avg.
113
 
114
+ `imatrix.dat` is in the repo if you want to re-quantize to IQ2_S, Q4_K_S, or anything else without redoing the calibration pass.
115
 
116
  ---
117
 
118
+ base model Β© Nex AGI Β· base architecture Β© Qwen team Β· llama.cpp Β© ggml-org. apache 2.0, same as the base.