IsValorum commited on
Commit
ce4f086
Β·
verified Β·
1 Parent(s): 2addcc3

Restore independent benchmark references; neutralize quantizer attribution

Browse files
Files changed (1) hide show
  1. README.md +25 -0
README.md CHANGED
@@ -59,6 +59,7 @@ quantized_by: IsValorum
59
  ## <a id="quick-navigation"></a>Quick Navigation Index
60
  - [Model Files & Technical Specifications](#model-specifications)
61
  - [Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)](#comparative-analysis)
 
62
  - [Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)](#laptop-benchmarks)
63
  - [The 24GB Miracle: Full 256K Context Runs In VRAM!](#context-scaling)
64
  - [Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50)](#throughput-projections)
@@ -66,6 +67,29 @@ quantized_by: IsValorum
66
  - [Model Inherent Behavior vs. Quantization Fidelity Notice](#quantization-fidelity)
67
  ---
68
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
69
  <a id="model-specifications"></a>
70
  ## Model Files & Technical Specifications
71
 
@@ -144,6 +168,7 @@ When running with full GPU offload (`-ngl 99`), KAT-Coder's fine-grained MoE arc
144
 
145
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Engineering Highlights |
146
  | :--- | :--- | :---: | :---: | :--- |
 
147
  | **NVIDIA RTX 4090 (24GB GDDR6X)** | Full GPU (`-ngl 99`) | **80 – 105+ tok/s** | **1,800 – 2,600+ tok/s** | Near-instantaneous code completion & refactoring |
148
  | **NVIDIA RTX 3090 (24GB GDDR6)** | Full GPU (`-ngl 99`) | **65 – 80+ tok/s** | **1,400 – 2,000+ tok/s** | Full 256k repository context in dedicated VRAM |
149
  | **NVIDIA RTX 4080 / 5070 (16GB)** | Partial offload (approx. 30 layers) | **35 – 45+ tok/s** | **800 – 1,200+ tok/s** | High-efficiency local coding assistant |
 
59
  ## <a id="quick-navigation"></a>Quick Navigation Index
60
  - [Model Files & Technical Specifications](#model-specifications)
61
  - [Comparative Quantization Analysis (vs. Flat Quants & Generic APEX)](#comparative-analysis)
62
+ - [Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)](#independent-benchmark)
63
  - [Everyday Laptop Benchmarks (DDR4 / DDR5 RAM)](#laptop-benchmarks)
64
  - [The 24GB Miracle: Full 256K Context Runs In VRAM!](#context-scaling)
65
  - [Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50)](#throughput-projections)
 
67
  - [Model Inherent Behavior vs. Quantization Fidelity Notice](#quantization-fidelity)
68
  ---
69
 
70
+ <a id="independent-benchmark"></a>
71
+ ### πŸ… Independent Benchmark of the APEX-I-MiniPlus Family (Occamy V2 Reference)
72
+
73
+
74
+ > [!NOTE]
75
+ > **External report:** [zephel01 independently benchmarked Occamy V2](https://note.com/zephel01/n/n71d3d7e6b70c?hl=en). The benchmark below was performed on **Occamy-1.0 APEX-I-MiniPlus V2**, not on this specific model. It is included as independent evidence of the broader MiniPlus quantization approach.
76
+
77
+ The **APEX-I-MiniPlus** quantization architecture powering this model was subjected to an extensive independent evaluation by Japanese AI researcher and evaluator [zephel01 (CoolZero)](https://note.com/zephel01/n/n71d3d7e6b70c?hl=en) on an **NVIDIA RTX 5090 (32GB)** workstation running `llama.cpp` CUDA `b11027` with FlashAttention (`-fa on -ctk q8_0 -ctv q8_0 -ngl 99`).
78
+
79
+ The evaluation tested the APEX-I hybrid MoE engine across **348 unseeded trials** on SWE-bench style multi-file Python bug-fixing tasks with hidden `pytest` suites (`llmbench`):
80
+
81
+ - **L6 Multi-File Code Generation (60 tasks):**
82
+ - **Context 32,768 (32K):** **93.3% Resolved** (46/60 tasks passed 5/5 consecutive trials; 20/20 on Easy–Hard).
83
+ - **Context 65,536 (65K):** **90.0% Resolved** (45/60 tasks passed 5/5 consecutive trials).
84
+ - **Match with 25–28 GB Models:** Matches or exceeds the resolution rate of full 25–28 GB models (such as `Ornith-1.5` and `Tiel-Coder` 35B-A3B) while consuming **over 10 GB less VRAM** (14.6 GB vs approx. 26 GB).
85
+ - **Extreme Context VRAM Scaling (The Hybrid DeltaNet SSM Advantage):**
86
+ - **32K Context:** **14.6 GB** total VRAM allocation.
87
+ - **65K Context:** **15.1 GB** total VRAM allocation (only **+0.5 GB VRAM** added when doubling context!).
88
+ - *Architectural Explanation:* Because 30 of the 40 layers utilize Linear Attention / DeltaNet SSM ($O(1)$ constant recurrence memory), only the 10 full-attention anchor layers expand the KV cache. This proves empirically that **65,536 context runs 100% in VRAM on consumer 16GB GPUs (RTX 4080 / RTX 5080)** without offloading to system RAM.
89
+ - **Measured Real-World Throughput:** Sustained single-stream generation of **approx. 247 – 251 tok/s** on NVIDIA RTX 5090.
90
+
91
+ ---
92
+
93
  <a id="model-specifications"></a>
94
  ## Model Files & Technical Specifications
95
 
 
168
 
169
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Engineering Highlights |
170
  | :--- | :--- | :---: | :---: | :--- |
171
+ | **NVIDIA RTX 5080 / 5090 (Blackwell)** | Full GPU (`-ngl 99`) | **approx. 247 – 251 tok/s** | **2,800 – 3,900+ tok/s** | Empirically verified on RTX 5090 by zephel01 (Occamy V2 Reference) |
172
  | **NVIDIA RTX 4090 (24GB GDDR6X)** | Full GPU (`-ngl 99`) | **80 – 105+ tok/s** | **1,800 – 2,600+ tok/s** | Near-instantaneous code completion & refactoring |
173
  | **NVIDIA RTX 3090 (24GB GDDR6)** | Full GPU (`-ngl 99`) | **65 – 80+ tok/s** | **1,400 – 2,000+ tok/s** | Full 256k repository context in dedicated VRAM |
174
  | **NVIDIA RTX 4080 / 5070 (16GB)** | Partial offload (approx. 30 layers) | **35 – 45+ tok/s** | **800 – 1,200+ tok/s** | High-efficiency local coding assistant |