IsValorum commited on
Commit
65dcb68
·
verified ·
1 Parent(s): e2150ee

Remove WikiText-2 description from V1 recommendation

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -42,7 +42,7 @@ quantized_by: IsValorum
42
  > - **MiniPlus V2.1 (System RAM Streaming Specialist with Deep Context):** Specially prepared to run **totally or partially in system RAM (DDR4/DDR5)** across massive multimodal and agentic context windows (up to 256k tokens). Upgrades all 40 shared foundation experts to `Q5_K`, armors attention gates in `Q8_0`, and uses linear CPU-friendly vectorization that eliminates AVX2 lookup stalls (+24 to 28+ tok/s). Depending on your processor and memory bandwidth (DDR4/DDR5), **streaming generation in system RAM can be almost as fast as having everything in VRAM**, while supporting deep context keeping the dedicated `Q8_0` multimodal vision projector (`mmproj`) explicitly loaded in GPU VRAM for instant screen parsing and OCR. All for **only approx. 180 MB more** (approx. 13.74 GiB vs approx. 13.56 GiB)—an overhead that is completely negligible in system RAM.
43
  >
44
  > **Which one should you choose? (Official Recommendation: V2.1)**
45
- > - **⭐ PRIMARY RECOMMENDATION — [Nex-N2.5-mini APEX-I-MiniPlus V2.1](https://huggingface.co/IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF):** For virtually all users and deployments, **V2.1 is the strictly recommended release**. Empirically verified on WikiText-2, V2.1 achieves an outstanding **Perplexity of 6.4725 ± 0.1635** (ΔPPL ≈ +0.07 from unquantized baseline (approx. 6.40)), matching the token fidelity of **Q5_K / Q6_K** class quantizations while weighing only **approx. 14.7 GB** (same footprint as Q3_K_M). Furthermore, it completely eliminates AVX2 CPU stalls, providing blistering **+24 to 28+ tok/s streaming** under system RAM offload.
46
  > - **MiniPlus V1 Legacy:** Maintained for architectural transparency and users seeking specialized configurations for their workflow.
47
  >
48
  > *Both editions are handcrafted and vastly outperform flat 3-bit quants and generic community APEX-I-Mini releases.*
 
42
  > - **MiniPlus V2.1 (System RAM Streaming Specialist with Deep Context):** Specially prepared to run **totally or partially in system RAM (DDR4/DDR5)** across massive multimodal and agentic context windows (up to 256k tokens). Upgrades all 40 shared foundation experts to `Q5_K`, armors attention gates in `Q8_0`, and uses linear CPU-friendly vectorization that eliminates AVX2 lookup stalls (+24 to 28+ tok/s). Depending on your processor and memory bandwidth (DDR4/DDR5), **streaming generation in system RAM can be almost as fast as having everything in VRAM**, while supporting deep context keeping the dedicated `Q8_0` multimodal vision projector (`mmproj`) explicitly loaded in GPU VRAM for instant screen parsing and OCR. All for **only approx. 180 MB more** (approx. 13.74 GiB vs approx. 13.56 GiB)—an overhead that is completely negligible in system RAM.
43
  >
44
  > **Which one should you choose? (Official Recommendation: V2.1)**
45
+ > - **⭐ PRIMARY RECOMMENDATION — [Nex-N2.5-mini APEX-I-MiniPlus V2.1](https://huggingface.co/IsValorum/Nex-N2.5-mini-APEX-I-MiniPlus-V2.1-GGUF):** For virtually all users and deployments, **V2.1 is the strictly recommended release**. V2.1 reports a **Perplexity of 6.4725 ± 0.1635** (ΔPPL ≈ +0.07 from unquantized baseline (approx. 6.40)), matching the token fidelity of **Q5_K / Q6_K** class quantizations while weighing only **approx. 14.7 GB** (same footprint as Q3_K_M). Furthermore, it completely eliminates AVX2 CPU stalls, providing blistering **+24 to 28+ tok/s streaming** under system RAM offload.
46
  > - **MiniPlus V1 Legacy:** Maintained for architectural transparency and users seeking specialized configurations for their workflow.
47
  >
48
  > *Both editions are handcrafted and vastly outperform flat 3-bit quants and generic community APEX-I-Mini releases.*