IsValorum commited on
Commit
d3fb3ca
·
verified ·
1 Parent(s): 0456a4c

Complete navigation index and remove duplicate support block

Browse files
Files changed (1) hide show
  1. README.md +20 -11
README.md CHANGED
@@ -34,18 +34,17 @@ pipeline_tag: image-text-to-text
34
  quantized_by: IsValorum
35
  ---
36
 
37
-
38
  ## <a id="quick-navigation"></a>Quick Navigation Index
39
- - [Model Files & Technical Specifications](#model-specifications)
40
- - [Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)](#comparative-analysis)
41
- - [Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)](#independent-benchmark)
42
- - [Bundled Q8_0 High-Precision Vision Projector](#vision-projector)
43
- - [Everyday Laptop Guidance (DDR4 / DDR5 RAM)](#laptop-benchmarks)
44
- - [The 24GB Miracle: Full 256K Context Runs In VRAM!](#context-scaling)
45
- - [Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50)](#throughput-projections)
46
- - [Recommended Generation Parameters (Creator Official)](#generation-parameters)
47
- - [Model Inherent Behavior vs. Quantization Fidelity Notice](#quantization-fidelity)
48
- - [Optional Support](#optional-support)
49
 
50
  > [!NOTE]
51
  > ### ARCHITECTURE SELECTION GUIDE — MINIPLUS V1 & V2.1 EDITIONS
@@ -92,6 +91,7 @@ quantized_by: IsValorum
92
  ---
93
 
94
  <a id="independent-benchmark"></a>
 
95
  ### 🏅 Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)
96
 
97
 
@@ -115,6 +115,7 @@ The evaluation tested the APEX-I hybrid MoE engine across **348 unseeded trials*
115
  ---
116
 
117
  <a id="model-specifications"></a>
 
118
  ## Model Files & Specifications
119
 
120
  | File Name | File Size | Memory Footprint | BPW | Description |
@@ -128,6 +129,7 @@ The evaluation tested the APEX-I hybrid MoE engine across **348 unseeded trials*
128
  ---
129
 
130
  <a id="comparative-analysis"></a>
 
131
  ## Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)
132
 
133
  Also, don't confuse **APEX-I-MiniPlus (Standard)** with a generic baseline APEX-I-Mini. Traditional APEX-I-Mini drops core experts aggressively to 2-bit `IQ2_S` and leaves `output.weight` at 3-bit `Q3_K_M`, which creates a noticeable perplexity hit on complex reasoning tasks. Standard MiniPlus avoids that degradation floor while keeping boundary layers in linear `Q3_K` for single-cycle vectorized AVX2 CPU dequantization (optimized for DDR4/DDR5 laptop streaming), while protecting output in `Q6_K` and routers in `F32`.
@@ -151,6 +153,7 @@ Take a look at the tensor-by-tensor comparison table below to inspect the exact
151
  ---
152
 
153
  <a id="vision-projector"></a>
 
154
  ## Bundled Q8_0 High-Precision Vision Projector
155
 
156
  Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual grounding, and multi-modal interaction.
@@ -160,7 +163,9 @@ Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual ground
160
  ---
161
 
162
  <a id="laptop-benchmarks"></a>
 
163
  ## Everyday Laptop Guidance
 
164
  ### *Empirically Verified in Unsloth Studio*
165
 
166
  - **GPU VRAM Offload:** Uses only **3.8 GB VRAM** (fits effortlessly on budget 4GB and 6GB laptop GPUs like the RTX 4050, 3050, or older 1660 Ti/2060).
@@ -171,6 +176,7 @@ Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual ground
171
  ---
172
 
173
  <a id="context-scaling"></a>
 
174
  ## The 24GB Miracle: Full 256K Context Runs In VRAM!
175
 
176
  | Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | **Total GPU VRAM (Est.)** | Hardware Verdict |
@@ -183,6 +189,7 @@ Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual ground
183
  ---
184
 
185
  <a id="throughput-projections"></a>
 
186
  ## Hardware Throughput Projections (RTX 30 / 40 / 50)
187
 
188
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
@@ -193,6 +200,7 @@ Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual ground
193
  | **Consumer Laptop (4GB GPU + 32GB RAM)**| Hybrid Offload | Hardware-dependent | Hardware-dependent | Smooth streaming from system DDR4/DDR5 RAM |
194
 
195
  <a id="generation-parameters"></a>
 
196
  ### ⚙️ Recommended Generation Parameters (nex-agi Official)
197
 
198
  Official sampling configuration recommended by [nex-agi](https://huggingface.co/nex-agi/Nex-N2.5-mini) for optimal generation quality across coding, browser-use, and agent evaluations:
@@ -210,6 +218,7 @@ Official sampling configuration recommended by [nex-agi](https://huggingface.co/
210
  > Any behavioral nuances, stylistic tendencies, domain-specific habits, or zero-shot edge-case oversights **stem entirely from the original unquantized checkpoint weights and fine-tuning distribution, NOT from the APEX-I quantization process.**
211
  > Handcrafted APEX-I-MiniPlus strictly preserves mathematical tensor fidelity—keeping 100% of expert routing matrices (`gate_inp`) in uncompressed `F32` (zero router drift), armoring the token output head in `Q6_K`, and safeguarding attention gates in `Q8_0`. Empirical verification confirms near-zero perplexity loss (ΔPPL ≈ +0.07), ensuring that token logits, routing decisions, and reasoning trajectories are mathematically faithful to the original base model.
212
 
 
213
  ## Optional Support
214
 
215
  > [!NOTE]
 
34
  quantized_by: IsValorum
35
  ---
36
 
 
37
  ## <a id="quick-navigation"></a>Quick Navigation Index
38
+ - [🏅 Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)](#toc-01)
39
+ - [Model Files & Specifications](#toc-02)
40
+ - [Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)](#toc-03)
41
+ - [Bundled Q8_0 High-Precision Vision Projector](#toc-04)
42
+ - [Everyday Laptop Guidance](#toc-05)
43
+ - [Empirically Verified in Unsloth Studio](#toc-06)
44
+ - [The 24GB Miracle: Full 256K Context Runs In VRAM!](#toc-07)
45
+ - [Hardware Throughput Projections (RTX 30 / 40 / 50)](#toc-08)
46
+ - [⚙️ Recommended Generation Parameters (nex-agi Official)](#toc-09)
47
+ - [Optional Support](#toc-10)
48
 
49
  > [!NOTE]
50
  > ### ARCHITECTURE SELECTION GUIDE — MINIPLUS V1 & V2.1 EDITIONS
 
91
  ---
92
 
93
  <a id="independent-benchmark"></a>
94
+ <a id="toc-01"></a>
95
  ### 🏅 Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)
96
 
97
 
 
115
  ---
116
 
117
  <a id="model-specifications"></a>
118
+ <a id="toc-02"></a>
119
  ## Model Files & Specifications
120
 
121
  | File Name | File Size | Memory Footprint | BPW | Description |
 
129
  ---
130
 
131
  <a id="comparative-analysis"></a>
132
+ <a id="toc-03"></a>
133
  ## Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)
134
 
135
  Also, don't confuse **APEX-I-MiniPlus (Standard)** with a generic baseline APEX-I-Mini. Traditional APEX-I-Mini drops core experts aggressively to 2-bit `IQ2_S` and leaves `output.weight` at 3-bit `Q3_K_M`, which creates a noticeable perplexity hit on complex reasoning tasks. Standard MiniPlus avoids that degradation floor while keeping boundary layers in linear `Q3_K` for single-cycle vectorized AVX2 CPU dequantization (optimized for DDR4/DDR5 laptop streaming), while protecting output in `Q6_K` and routers in `F32`.
 
153
  ---
154
 
155
  <a id="vision-projector"></a>
156
+ <a id="toc-04"></a>
157
  ## Bundled Q8_0 High-Precision Vision Projector
158
 
159
  Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual grounding, and multi-modal interaction.
 
163
  ---
164
 
165
  <a id="laptop-benchmarks"></a>
166
+ <a id="toc-05"></a>
167
  ## Everyday Laptop Guidance
168
+ <a id="toc-06"></a>
169
  ### *Empirically Verified in Unsloth Studio*
170
 
171
  - **GPU VRAM Offload:** Uses only **3.8 GB VRAM** (fits effortlessly on budget 4GB and 6GB laptop GPUs like the RTX 4050, 3050, or older 1660 Ti/2060).
 
176
  ---
177
 
178
  <a id="context-scaling"></a>
179
+ <a id="toc-07"></a>
180
  ## The 24GB Miracle: Full 256K Context Runs In VRAM!
181
 
182
  | Context Length | Model Weights (Est.) | KV Cache (q8_0, 4 slots) | Compute Buffers | **Total GPU VRAM (Est.)** | Hardware Verdict |
 
189
  ---
190
 
191
  <a id="throughput-projections"></a>
192
+ <a id="toc-08"></a>
193
  ## Hardware Throughput Projections (RTX 30 / 40 / 50)
194
 
195
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
 
200
  | **Consumer Laptop (4GB GPU + 32GB RAM)**| Hybrid Offload | Hardware-dependent | Hardware-dependent | Smooth streaming from system DDR4/DDR5 RAM |
201
 
202
  <a id="generation-parameters"></a>
203
+ <a id="toc-09"></a>
204
  ### ⚙️ Recommended Generation Parameters (nex-agi Official)
205
 
206
  Official sampling configuration recommended by [nex-agi](https://huggingface.co/nex-agi/Nex-N2.5-mini) for optimal generation quality across coding, browser-use, and agent evaluations:
 
218
  > Any behavioral nuances, stylistic tendencies, domain-specific habits, or zero-shot edge-case oversights **stem entirely from the original unquantized checkpoint weights and fine-tuning distribution, NOT from the APEX-I quantization process.**
219
  > Handcrafted APEX-I-MiniPlus strictly preserves mathematical tensor fidelity—keeping 100% of expert routing matrices (`gate_inp`) in uncompressed `F32` (zero router drift), armoring the token output head in `Q6_K`, and safeguarding attention gates in `Q8_0`. Empirical verification confirms near-zero perplexity loss (ΔPPL ≈ +0.07), ensuring that token logits, routing decisions, and reasoning trajectories are mathematically faithful to the original base model.
220
 
221
+ <a id="toc-10"></a>
222
  ## Optional Support
223
 
224
  > [!NOTE]