IsValorum commited on
Commit
7cac7fb
Β·
verified Β·
1 Parent(s): 1cd0a4d

Clean duplicate banners, remove conflicting legacy text, and restore accurate Q8_0 specs

Browse files
Files changed (1) hide show
  1. README.md +2 -13
README.md CHANGED
@@ -38,8 +38,8 @@ quantized_by: IsValorum
38
  > ### πŸ›οΈ ARCHITECTURE SELECTION GUIDE β€” MINIPLUS V1 & V2.1 EDITIONS
39
  > This repository hosts the **MiniPlus V1** edition of **Nex-N2.5-mini**. Our releases are precision-engineered for specific hardware budgets and memory topologies. **V1 is NOT obsolete or inferior; it represents our leanest, most agile operating profile:**
40
  >
41
- > - **MiniPlus V1 (Lean & Agile Profile):** Highly compact footprint with uncompressed `F32` router gates, a fully armored `Q6_K` output head, and `IQ3_XXS` core experts. **Both V1 and V2.1 run flawlessly with the vast majority of the model residing in system RAM (DDR4/DDR5)**, thanks to linear CPU-friendly vectorization that avoids lookup stalls. V1 is dramatically superior to generic community APEX-I-Mini releases (which crush core reasoning down to 2-bit `IQ2_S`) and flat 3-bit quants.
42
- > - **MiniPlus V2.1 (Enhanced Protection & Refined Throughput):** Upgrades all 40 shared foundation experts to high-precision `Q5_K` and armors full attention anchor layers (`Q4_K`/`Q6_K`). This improves inference speed and stability even further for **only a few hundred megabytes extra** (~13.74 GiB vs ~13.56 GiB)β€”an overhead that is completely negligible when running in system RAM (16GB/32GB/64GB).
43
  >
44
  > πŸ’‘ **Which one should you choose?**
45
  > - **If your system has strict memory constraints:** **V1** delivers uncompromising reasoning at our lowest RAM footprint.
@@ -49,17 +49,6 @@ quantized_by: IsValorum
49
  > πŸ‘‰ To explore or download the **V2.1** edition of Nex-N2.5-mini, visit:
50
  > **[IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF](https://huggingface.co/IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF)**
51
 
52
- # πŸ›οΈ ARCHITECTURE SELECTION GUIDE β€” MINIPLUS TIER OVERVIEW
53
- > Every edition of the **MiniPlus** family is a precision-engineered, handcrafted quantization designed for specific hardware constraints and memory footprints. **None of these releases are obsolete; each represents an optimal operating point tailored to your system budget:**
54
- >
55
- > - **MiniPlus V1 (Lean & Agile Foundation):** Maximum compactness and ultra-fast throughput with minimal RAM/VRAM footprint. Even in this lightest profile, **V1 dramatically outperforms generic community APEX-I-Mini releases** (which aggressively downgrade core reasoning to flat 2-bit `IQ2_S` and leave attention and output heads degraded). V1 provides uncompressed `F32` router gates, `Q6_K` output head protection, and `IQ3_XXS` core experts.
56
- > - **MiniPlus V2 (Expanded Edge Defense):** Adds wider protective envelopes on edge layers (10 layers in `IQ3_S` + `IQ4_NL` shared experts + `Q8_0` attention gates) for workstations with an extra ~1 GB of headroom seeking enhanced attention stability.
57
- > - **MiniPlus V2.1 (Current Long-Context Standard):** Upgrades shared experts to `Q5_K` across all 40 layers and optimizes full attention tensors (`Q4_K`/`Q6_K`) to eliminate AVX2 CPU dequantization stalls during deep offloading.
58
- >
59
- > πŸ’‘ *Choose the version that fits your exact hardware budget! All editions provide rock-solid reasoning and far exceed generic community quants.*
60
- > πŸ‘‰ If your workstation has sufficient memory headroom and you want the latest V2.1 specification, you can find it directly at:
61
- > **[IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF](https://huggingface.co/IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF)**
62
-
63
  ---
64
 
65
  ## <a id="quick-navigation"></a>⚑ Quick Navigation Index
 
38
  > ### πŸ›οΈ ARCHITECTURE SELECTION GUIDE β€” MINIPLUS V1 & V2.1 EDITIONS
39
  > This repository hosts the **MiniPlus V1** edition of **Nex-N2.5-mini**. Our releases are precision-engineered for specific hardware budgets and memory topologies. **V1 is NOT obsolete or inferior; it represents our leanest, most agile operating profile:**
40
  >
41
+ > - **MiniPlus V1 (Lean & Agile Profile):** Highly compact footprint with uncompressed `F32` router gates, a fully armored `Q6_K` output head, `Q8_0` attention gates, and `IQ3_XXS` core experts. **Both V1 and V2.1 run flawlessly with the vast majority of the model residing in system RAM (DDR4/DDR5)**, thanks to linear CPU-friendly vectorization that avoids lookup stalls. V1 is dramatically superior to generic community APEX-I-Mini releases (which crush core reasoning down to 2-bit `IQ2_S`) and flat 3-bit quants.
42
+ > - **MiniPlus V2.1 (Enhanced Protection & Refined Throughput):** Upgrades all 40 shared foundation experts to high-precision `Q5_K` and armors full attention anchor layers (`Q4_K`/`Q6_K`). This improves inference speed and stability even further for **only ~180 MB more** (~13.74 GiB vs ~13.56 GiB)β€”an overhead that is completely negligible when running in system RAM (16GB/32GB/64GB).
43
  >
44
  > πŸ’‘ **Which one should you choose?**
45
  > - **If your system has strict memory constraints:** **V1** delivers uncompromising reasoning at our lowest RAM footprint.
 
49
  > πŸ‘‰ To explore or download the **V2.1** edition of Nex-N2.5-mini, visit:
50
  > **[IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF](https://huggingface.co/IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF)**
51
 
 
 
 
 
 
 
 
 
 
 
 
52
  ---
53
 
54
  ## <a id="quick-navigation"></a>⚑ Quick Navigation Index