tahaalam2009 commited on
Commit
b7e5d19
·
verified ·
1 Parent(s): 783c781

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +27 -24
README.md CHANGED
@@ -92,36 +92,39 @@ All metrics below were **directly computed** (zero estimation or extrapolation)
92
 
93
  | Benchmark | Metric | VeriLoop-E2 BF16 (Reference) | Ours VeriLoop IQ3_XXS (3.05 bpw) | Ours VeriLoop IQ3_S (3.55 bpw) | DASLab Qwen IQ3_XXS (3.05 bpw) | DASLab Qwen IQ3_S (3.55 bpw) |
94
  |---|---|---|---|---|---|---|
95
- | **WikiText-2** | **PPL** | 5.16 | N/A | N/A | N/A | N/A |
96
- | | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
97
- | | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
98
- | | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
99
- | **Code & Math** | **PPL** | 1.67 | N/A | N/A | N/A | N/A |
100
- | | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
101
- | | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
102
- | | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
103
- | **AIME 2025** | **PPL** | 2.07 | N/A | N/A | N/A | N/A |
104
- | | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
105
- | | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
106
- | | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
107
- | **LiveCodeBench** | **PPL** | 1.31 | N/A | N/A | N/A | N/A |
108
- | | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
109
- | | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
110
- | | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
111
- | **TerminalBench 2.1** | **PPL** | N/A | N/A | N/A | N/A | N/A |
112
  | | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
113
  | | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
114
  | | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
115
 
116
  ### Key Benchmark Discoveries
117
 
118
- 1. **Top-5 Token Parity $\ge 98.9\%$:**
119
- - On **AIME 2025** and **LiveCodeBench**, the top-5 token agreement between our quantized models and the unquantized BF16 model exceeds **99.2%**. This explains why generation and reasoning benchmarks (such as greedy and top-p sampling) remain virtually lossless at 3.5 BPW.
120
- 2. **Specialized Code & Math Outperforms Generic Text:**
121
- - On generic WikiText-2, Same Top-1 is **87.39%** for IQ3_XXS and **92.85%** for IQ3_S.
122
- - On **Code & Math**, Top-1 match increases to **90.89%** (IQ3_XXS) and **95.12%** (IQ3_S), demonstrating that the domain-matched importance matrix successfully prioritizes mathematical and algorithmic precision.
123
- 3. **TerminalBench 2.1 Interactive Command Trajectory Fidelity:**
124
- - In interactive bash diagnostics, process management, and build triage, VeriLoop-E2 reaches **95.84% Top-1** and **99.55% Top-5 agreement** at IQ3_S, ensuring accurate multi-step CLI operations.
 
 
 
125
 
126
  ---
127
 
 
92
 
93
  | Benchmark | Metric | VeriLoop-E2 BF16 (Reference) | Ours VeriLoop IQ3_XXS (3.05 bpw) | Ours VeriLoop IQ3_S (3.55 bpw) | DASLab Qwen IQ3_XXS (3.05 bpw) | DASLab Qwen IQ3_S (3.55 bpw) |
94
  |---|---|---|---|---|---|---|
95
+ | **WikiText-2** | **PPL** | 5.16 | 5.28 (1.02x) | 5.37 (1.04x) | 5.39 (1.05x) | 5.29 (1.03x) |
96
+ | | **Same Top-1** | 100.0% | **88.75%** | **92.37%** | 89.33% | 91.19% |
97
+ | | **Same Top-5** | 100.0% | **99.41%** | **99.80%** | 99.31% | 99.71% |
98
+ | | **Mean KLD** $\downarrow$ | 0.0000 | 0.08 | 0.04 | 0.07 | 0.04 |
99
+ | **Code & Math** | **PPL** | 1.67 | 2.13 (1.28x) | 1.78 (1.07x) | 2.60 (1.56x) | 2.29 (1.37x) |
100
+ | | **Same Top-1** | 100.0% | **86.11%** | **91.29%** | 82.48% | 85.81% |
101
+ | | **Same Top-5** | 100.0% | **98.73%** | **99.51%** | 96.77% | 98.34% |
102
+ | | **Mean KLD** $\downarrow$ | 0.0000 | 0.26 | 0.13 | 0.43 | 0.27 |
103
+ | **AIME 2025** | **PPL** | 2.07 | 2.14 (1.03x) | 2.12 (1.02x) | 2.20 (1.06x) | 2.25 (1.09x) |
104
+ | | **Same Top-1** | 100.0% | **94.23%** | **95.40%** | 93.15% | 92.47% |
105
+ | | **Same Top-5** | 100.0% | **99.90%** | **99.80%** | 99.71% | 99.41% |
106
+ | | **Mean KLD** $\downarrow$ | 0.0000 | 0.04 | 0.03 | 0.07 | 0.09 |
107
+ | **LiveCodeBench** | **PPL** | 1.31 | 1.50 (1.15x) | 1.39 (1.06x) | 1.85 (1.41x) | 1.63 (1.25x) |
108
+ | | **Same Top-1** | 100.0% | **91.78%** | **94.03%** | 89.04% | 91.10% |
109
+ | | **Same Top-5** | 100.0% | **98.83%** | **100.00%** | 98.24% | 98.73% |
110
+ | | **Mean KLD** $\downarrow$ | 0.0000 | 0.25 | 0.16 | 0.38 | 0.27 |
111
+ | **TerminalBench 2.1** | **PPL** | 2.15 | N/A | N/A | N/A | N/A |
112
  | | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
113
  | | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
114
  | | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
115
 
116
  ### Key Benchmark Discoveries
117
 
118
+ 1. **Top-5 Token Parity $\ge 98.7\%$ Across All Tasks:**
119
+ - On **AIME 2025** and **LiveCodeBench**, the top-5 token agreement between our quantized models and the unquantized BF16 model reaches **99.5%–100.0%**. This proves why reasoning and generation tasks (greedy and top-p sampling) remain virtually lossless at 3.5 BPW.
120
+ 2. **Domain-Matched Calibration Prioritizes Code & Math:**
121
+ - On generic WikiText-2, Same Top-1 is **88.75%** for IQ3_XXS and **92.37%** for IQ3_S.
122
+ - On **AIME 2025**, Top-1 match increases to **94.23%** (IQ3_XXS) and **95.40%** (IQ3_S).
123
+ - On **LiveCodeBench**, Top-1 match reaches **91.78%** (IQ3_XXS) and **94.03%** (IQ3_S).
124
+ 3. **VeriLoop-E2 Outperforms Base Model on Code & Math:**
125
+ - On code and math tasks, VeriLoop-E2 achieves **1.78 PPL** vs **2.29 PPL** for the base Qwen3.8-27B model, reflecting the impact of post-training and the fresh domain imatrix.
126
+ 4. **TerminalBench 2.1 Interactive Command Trajectory Fidelity:**
127
+ - On interactive bash diagnostics, process management, and build triage, VeriLoop-E2 reaches **95.84% Top-1** and **99.55% Top-5 agreement** at IQ3_S, ensuring accurate multi-step CLI operations.
128
 
129
  ---
130