IsValorum commited on
Commit
a2869d9
Β·
verified Β·
1 Parent(s): 5a4b914

Separate author sampling recommendation from zephel01 llmbench setting & explicit Occamy V2 references

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -74,7 +74,7 @@ quantized_by: IsValorum
74
 
75
 
76
  > [!NOTE]
77
- > The benchmark below was performed on **Occamy-1.0 APEX-I-MiniPlus V2**, not on this specific model. It is included as independent evidence of the broader MiniPlus quantization approach.
78
 
79
  The **APEX-I-MiniPlus** quantization architecture powering this model was subjected to an extensive independent evaluation by Japanese AI researcher and evaluator [zephel01 (CoolZero)](https://note.com/zephel01/n/n71d3d7e6b70c?hl=en) on an **NVIDIA RTX 5090 (32GB)** workstation running `llama.cpp` CUDA `b11027` with FlashAttention (`-fa on -ctk q8_0 -ctv q8_0 -ngl 99`).
80
 
@@ -165,7 +165,7 @@ Unlike text-only MoEs, Nex-N2.5-mini is designed for computer use, visual ground
165
 
166
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
167
  | :--- | :--- | :---: | :---: | : |
168
- | **NVIDIA RTX 5080 / 5090 (Blackwell)** | Full GPU (`-ngl 99`) | **approx. 247 – 251 tok/s** | **2,800 – 3,900+ tok/s** | Empirically verified on RTX 5090 by zephel01 |
169
  | **NVIDIA RTX 4090 (24GB GDDR6X)** | Full GPU (`-ngl 99`) + mmproj | **75 – 100+ tok/s** | **1,700 – 2,500+ tok/s** | Real-time computer-use screen analysis & tool calling |
170
  | **NVIDIA RTX 3090 (24GB GDDR6)** | Full GPU (`-ngl 99`) + mmproj | **62 – 78+ tok/s** | **1,350 – 1,950+ tok/s** | Full 256k multi-modal context in dedicated VRAM |
171
  | **Consumer Laptop (4GB GPU + 32GB RAM)**| Hybrid Offload | **20 – 24+ tok/s** | **300 – 420+ tok/s** | Smooth streaming from system DDR4/DDR5 RAM |
 
74
 
75
 
76
  > [!NOTE]
77
+ > **Occamy V2 Reference Notice:** The benchmark below was performed on **Occamy-1.0 APEX-I-MiniPlus V2**, not on this specific Nex-N2.5 model. It is included as independent evidence of the broader APEX-I-MiniPlus quantization approach and hybrid MoE architecture.
78
 
79
  The **APEX-I-MiniPlus** quantization architecture powering this model was subjected to an extensive independent evaluation by Japanese AI researcher and evaluator [zephel01 (CoolZero)](https://note.com/zephel01/n/n71d3d7e6b70c?hl=en) on an **NVIDIA RTX 5090 (32GB)** workstation running `llama.cpp` CUDA `b11027` with FlashAttention (`-fa on -ctk q8_0 -ctv q8_0 -ngl 99`).
80
 
 
165
 
166
  | Hardware Target | Offload Mode | Generation Speed (Est.) | Prompt Prefill Speed (Est.) | Highlights |
167
  | :--- | :--- | :---: | :---: | : |
168
+ | **NVIDIA RTX 5080 / 5090 (Blackwell)** | Full GPU (`-ngl 99`) | **approx. 247 – 251 tok/s** | **2,800 – 3,900+ tok/s** | Empirically verified on RTX 5090 by zephel01 (Occamy V2 Reference) |
169
  | **NVIDIA RTX 4090 (24GB GDDR6X)** | Full GPU (`-ngl 99`) + mmproj | **75 – 100+ tok/s** | **1,700 – 2,500+ tok/s** | Real-time computer-use screen analysis & tool calling |
170
  | **NVIDIA RTX 3090 (24GB GDDR6)** | Full GPU (`-ngl 99`) + mmproj | **62 – 78+ tok/s** | **1,350 – 1,950+ tok/s** | Full 256k multi-modal context in dedicated VRAM |
171
  | **Consumer Laptop (4GB GPU + 32GB RAM)**| Hybrid Offload | **20 – 24+ tok/s** | **300 – 420+ tok/s** | Smooth streaming from system DDR4/DDR5 RAM |