IsValorum commited on
Commit
cd023fe
·
verified ·
1 Parent(s): a0f439a

Restore independent benchmark references; neutralize quantizer attribution

Browse files
Files changed (1) hide show
  1. README.md +25 -0
README.md CHANGED
@@ -60,6 +60,7 @@ quantized_by: IsValorum
60
  ## <a id="quick-navigation"></a>Quick Navigation Index
61
  - [Model Files & Technical Specifications](#model-specifications)
62
  - [Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)](#comparative-analysis)
 
63
  - [Bundled Q8_0 High-Precision Vision Projector](#vision-projector)
64
  - [Everyday Laptop Guidance (DDR4 / DDR5 RAM)](#laptop-benchmarks)
65
  - [The 24GB Miracle: Full 256K Context Runs In VRAM!](#context-scaling)
@@ -68,6 +69,29 @@ quantized_by: IsValorum
68
  - [Model Inherent Behavior vs. Quantization Fidelity Notice](#quantization-fidelity)
69
  ---
70
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71
  <a id="model-specifications"></a>
72
  ## Model Files & Specifications
73
 
@@ -141,6 +165,7 @@ Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual ground
141
 
142
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
143
  | :--- | :--- | :---: | :---: | : |
 
144
  | **NVIDIA RTX 4090 (24GB GDDR6X)** | Full GPU (`-ngl 99`) + mmproj | **75 – 100+ tok/s** | **1,700 – 2,500+ tok/s** | Real-time computer-use screen analysis & tool calling |
145
  | **NVIDIA RTX 3090 (24GB GDDR6)** | Full GPU (`-ngl 99`) + mmproj | **62 – 78+ tok/s** | **1,350 – 1,950+ tok/s** | Full 256k multi-modal context in dedicated VRAM |
146
  | **Consumer Laptop (4GB GPU + 32GB RAM)**| Hybrid Offload | Hardware-dependent | Hardware-dependent | Smooth streaming from system DDR4/DDR5 RAM |
 
60
  ## <a id="quick-navigation"></a>Quick Navigation Index
61
  - [Model Files & Technical Specifications](#model-specifications)
62
  - [Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)](#comparative-analysis)
63
+ - [Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)](#independent-benchmark)
64
  - [Bundled Q8_0 High-Precision Vision Projector](#vision-projector)
65
  - [Everyday Laptop Guidance (DDR4 / DDR5 RAM)](#laptop-benchmarks)
66
  - [The 24GB Miracle: Full 256K Context Runs In VRAM!](#context-scaling)
 
69
  - [Model Inherent Behavior vs. Quantization Fidelity Notice](#quantization-fidelity)
70
  ---
71
 
72
+ <a id="independent-benchmark"></a>
73
+ ### 🏅 Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)
74
+
75
+
76
+ > [!NOTE]
77
+ > **Occamy V2 Reference Notice:** **External report:** [zephel01 independently benchmarked Occamy V2](https://note.com/zephel01/n/n71d3d7e6b70c?hl=en). The benchmark below was performed on **Occamy-1.0 APEX-I-MiniPlus V2**, not on this specific Nex-N2.5 model. It is included as independent evidence of the broader APEX-I-MiniPlus quantization approach and hybrid MoE architecture.
78
+
79
+ The **APEX-I-MiniPlus** quantization architecture powering this model was subjected to an extensive independent evaluation by Japanese AI researcher and evaluator [zephel01 (CoolZero)](https://note.com/zephel01/n/n71d3d7e6b70c?hl=en) on an **NVIDIA RTX 5090 (32GB)** workstation running `llama.cpp` CUDA `b11027` with FlashAttention (`-fa on -ctk q8_0 -ctv q8_0 -ngl 99`).
80
+
81
+ The evaluation tested the APEX-I hybrid MoE engine across **348 unseeded trials** on SWE-bench style multi-file Python bug-fixing tasks with hidden `pytest` suites (`llmbench`):
82
+
83
+ - **L6 Multi-File Code Generation (60 tasks):**
84
+ - **Context 32,768 (32K):** **93.3% Resolved** (46/60 tasks passed 5/5 consecutive trials; 20/20 on Easy–Hard).
85
+ - **Context 65,536 (65K):** **90.0% Resolved** (45/60 tasks passed 5/5 consecutive trials).
86
+ - **Match with 25–28 GB Models:** Matches or exceeds the resolution rate of full 25–28 GB models (such as `Ornith-1.5` and `Tiel-Coder` 35B-A3B) while consuming **over 10 GB less VRAM** (14.6 GB vs approx. 26 GB).
87
+ - **Extreme Context VRAM Scaling (The Hybrid DeltaNet SSM Advantage):**
88
+ - **32K Context:** **14.6 GB** total VRAM allocation.
89
+ - **65K Context:** **15.1 GB** total VRAM allocation (only **+0.5 GB VRAM** added when doubling context!).
90
+ - *Architectural Explanation:* Because 30 of the 40 layers utilize Linear Attention / DeltaNet SSM ($O(1)$ constant recurrence memory), only the 10 full-attention anchor layers expand the KV cache. This proves empirically that **65,536 context runs 100% in VRAM on consumer 16GB GPUs (RTX 4080 / RTX 5080)** without offloading to system RAM.
91
+ - **Measured Real-World Throughput:** Sustained single-stream generation of **approx. 247 – 251 tok/s** on NVIDIA RTX 5090.
92
+
93
+ ---
94
+
95
  <a id="model-specifications"></a>
96
  ## Model Files & Specifications
97
 
 
165
 
166
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
167
  | :--- | :--- | :---: | :---: | : |
168
+ | **NVIDIA RTX 5080 / 5090 (Blackwell)** | Full GPU (`-ngl 99`) | **approx. 247 – 251 tok/s** | **2,800 – 3,900+ tok/s** | Empirically verified on RTX 5090 by zephel01 (Occamy V2 Reference) |
169
  | **NVIDIA RTX 4090 (24GB GDDR6X)** | Full GPU (`-ngl 99`) + mmproj | **75 – 100+ tok/s** | **1,700 – 2,500+ tok/s** | Real-time computer-use screen analysis & tool calling |
170
  | **NVIDIA RTX 3090 (24GB GDDR6)** | Full GPU (`-ngl 99`) + mmproj | **62 – 78+ tok/s** | **1,350 – 1,950+ tok/s** | Full 256k multi-modal context in dedicated VRAM |
171
  | **Consumer Laptop (4GB GPU + 32GB RAM)**| Hybrid Offload | Hardware-dependent | Hardware-dependent | Smooth streaming from system DDR4/DDR5 RAM |