ConorWang commited on
Commit
b134bff
·
verified ·
1 Parent(s): 4763719

Upload 3 files

Browse files
Files changed (3) hide show
  1. QUANTIZATION_QUALITY.md +54 -74
  2. README.md +84 -185
  3. RELEASE_MANIFEST.json +67 -403
QUANTIZATION_QUALITY.md CHANGED
@@ -7,36 +7,34 @@
7
  **Runtime:** llama.cpp
8
  **Reference precision:** BF16
9
  **Measured ladder:** Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_M, IQ2_S, IQ1_M
10
- **Parent-model benchmark re-run per quant:** No
11
 
12
- > This document separates **model capability** from **quantization fidelity**. Parent-model benchmark scores establish the parent VeriLoop E2 capability. The measurements below quantify how closely each GGUF tier preserves the canonical BF16 reference under one frozen paired protocol.
13
 
14
  ## 1. Executive decision
15
 
16
- **Overall sweet spot — Q6_K.** **58.96%** smaller than BF16 with PPL parity within reported uncertainty, Mean KLD **0.004409 ± 0.000953**, Same top-p **98.204 ± 0.147%**.
17
 
18
- **Memory-quality sweet spot — Q5_K_M.** **62.16%** smaller than BF16 and **7.78%** smaller than Q6_K, with **+0.4450%** PPL drift, Mean KLD **0.006919 ± 0.000945**, Same top-p **97.251 ± 0.181%**.
19
 
20
- **Current released low-footprint sweet spot — IQ2_S.** **16.798508 GiB**, full release verification PASS, PPL ratio **1.003457 ± 0.002100**, Mean KLD **0.014023 ± 0.001408**, Same top-p **95.516 ± 0.229%**.
21
 
22
- **Ultra-low-footprint frontier candidate — IQ1_M.** The measured V2 artifact is **16.790078 GiB**, **66.4955% smaller than BF16**, and passes the unchanged hard and engineering-reserve fidelity gates with PPL ratio **1.003191 ± 0.002175**, Mean KLD **0.014357 ± 0.001317**, Same top-p **95.870 ± 0.220%**. It is smaller than IQ2_S and has better PPL/Same-top point estimates, but slightly worse KLD/RMS. The accepted rebuild reproduced the same policy and froze the local SHA256. Stock runtime, MTP, and remote verification still remain before IQ1_M is promoted to a released sweet spot.
23
-
24
- **Objective low-footprint conclusion.** IQ1_M and IQ2_S are neighboring Pareto points rather than a clean first/second ranking. IQ2_S currently wins on release maturity and Mean KLD; IQ1_M wins on measured footprint, PPL ratio, and Same top-p. Their uncertainty intervals substantially overlap.
25
 
26
  ## 2. Precision matrix
27
 
28
- | Tier | Role | Main size | Effective density | Mean PPL | Mean KLD | Same top-p | State |
29
- |---|---|---:|---:|---:|---:|---:|---|
30
- | BF16 | Reference | **50.112870 GiB** | 16-bit-class | **4.840423 ± 0.119931** | 0 | 100% | Verified reference |
31
- | Q8_0 | High fidelity | **26.631882 GiB** | **8.50 BPW** | **4.843536 ± 0.120062** | **0.002176 ± 0.000668** | **98.815 ± 0.120%** | PASS |
32
- | **Q6_K** | **Overall sweet spot** | **20.565961 GiB** | **6.57 BPW** | **4.838514 ± 0.119694** | **0.004409 ± 0.000953** | **98.204 ± 0.147%** | PASS |
33
- | **Q5_K_M** | **Memory-quality sweet spot** | **18.965142 GiB** | **6.05 BPW** | **4.861965 ± 0.120630** | **0.006919 ± 0.000945** | **97.251 ± 0.181%** | PASS |
34
- | Q4_K_M | Intermediate candidate | **18.301080 GiB** | **5.84 BPW** | **4.863760 ± 0.120720** | **0.009700 ± 0.001110** | **96.786 ± 0.195%** | Quantitative PASS |
35
- | Q3_K_M | Adjacent low-footprint release | **16.825745 GiB** | **5.37 BPW** | **4.860222 ± 0.120606** | **0.014349 ± 0.001985** | **95.919 ± 0.219%** | Full release PASS |
36
- | **IQ2_S** | **Released low-footprint sweet spot** | **16.798508 GiB** | **5.36 BPW** | **4.857159 ± 0.120482** | **0.014023 ± 0.001408** | **95.516 ± 0.229%** | **Full release PASS** |
37
- | **IQ1_M** | **Ultra-low-footprint frontier** | **16.790078 GiB** | **5.36 BPW** | **4.855870 ± 0.120480** | **0.014357 ± 0.001317** | **95.870 ± 0.220%** | **Hard + reserve quantitative PASS; final release gates pending** |
38
 
39
- ## 3. Performance-loss benchmark
40
 
41
  | Tier | PPL ratio | Relative PPL change | Mean KLD | Same top-p | Top-token disagreement | RMS Δp | log-PPL corr. |
42
  |---|---:|---:|---:|---:|---:|---:|---:|
@@ -48,7 +46,7 @@
48
  | IQ2_S | **1.003457 ± 0.002100** | **+0.3457%** | **0.014023 ± 0.001408** | **95.516 ± 0.229%** | **4.484 pp** | **3.679 ± 0.231%** | **99.64%** |
49
  | **IQ1_M** | **1.003191 ± 0.002175** | **+0.3191%** | **0.014357 ± 0.001317** | **95.870 ± 0.220%** | **4.130 pp** | **3.691 ± 0.205%** | **99.62%** |
50
 
51
- **Claim boundary.** These are **quantization-retention measurements**, not downstream task-score loss. The nine parent-model benchmarks were not rerun independently per quant. Therefore **+0.3191% PPL does not mean a 0.3191% loss on SWE-bench, Terminal-Bench, DeepSWE, AIME, GPQA, Apex, coding-agent, mathematics, or physics performance.**
52
 
53
  ## 4. Frozen BF16-paired protocol
54
 
@@ -65,7 +63,7 @@
65
  | Batch / micro-batch | 512 / 512 |
66
  | Evaluator | `llama-perplexity` |
67
  | Reference method | `--kl-divergence-base` |
68
- | Candidate method | `--kl-divergence` |
69
  | BF16 logits reused | Yes |
70
  | llama.cpp revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` |
71
 
@@ -81,7 +79,7 @@
81
  | Mean Δp | 0 | **−0.042 ± 0.041%** | small directional bias |
82
  | log-PPL correlation | 100% | **99.62%** | high agreement |
83
 
84
- ### IQ1_M KLD distribution
85
 
86
  | Statistic | Value |
87
  |---|---:|
@@ -95,11 +93,11 @@
95
 
96
  ## 6. IQ1_M exact tensor policy
97
 
98
- The measured Q1 V2 policy differs from the successful IQ2_S release by exactly **one tensor**:
99
 
100
  - `blk.1.ffn_down.weight`: **IQ2_S → IQ1_M**
101
- - `blk.0.ffn_down.weight`: stays **IQ2_S**
102
- - `blk.3.ffn_down.weight`: stays **IQ2_S**
103
 
104
  | Precision | Count |
105
  |---|---:|
@@ -111,7 +109,7 @@ The measured Q1 V2 policy differs from the successful IQ2_S release by exactly *
111
  | Q6_K | **2** |
112
  | **Total** | **851** |
113
 
114
- | Build item | Measured V2 value |
115
  |---|---|
116
  | Source | Canonical BF16 |
117
  | Low-bit requantization | **No** |
@@ -123,20 +121,21 @@ The measured Q1 V2 policy differs from the successful IQ2_S release by exactly *
123
  | End-to-end quantization timer | **168.26090 s** |
124
  | SHA256 | `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b` |
125
 
126
- ## 7. IQ1_M quantitative gates
127
-
128
- | Gate | Hard | Reserve | Measured V2 | Result |
129
- |---|---:|---:|---:|---|
130
- | PPL ratio | ≤1.0150 | ≤1.0135 | **1.003191** | **PASS / PASS** |
131
- | Mean KLD | ≤0.0150 | ≤0.0145 | **0.014357** | **PASS / PASS** |
132
- | Same top-p | ≥95.0% | ≥95.5% | **95.870%** | **PASS / PASS** |
133
- | Tensor structure | exact | exact | **851; 0 mismatch** | **PASS** |
134
 
135
- Reserve headroom is **0.010309** on PPL ratio, **0.000143** on Mean KLD, and **0.370 pp** on Same top-p. Mean KLD is the binding Q1 reserve metric.
 
 
 
 
 
 
 
 
 
 
136
 
137
- The V2 artifact was rejected by an additional experimental rule requiring Q1 to beat IQ2_S on **every selected point estimate**. That rule was stricter than the frozen release-quality gate and has been removed from V3. The **hard and reserve thresholds themselves are unchanged**.
138
-
139
- ## 8. IQ1_M vs IQ2_S vs Q3_K_M frontier
140
 
141
  | Metric | Q3_K_M | IQ2_S | IQ1_M |
142
  |---|---:|---:|---:|
@@ -146,47 +145,28 @@ The V2 artifact was rejected by an additional experimental rule requiring Q1 to
146
  | Same top-p | **95.919%** | **95.516%** | **95.870%** |
147
  | RMS Δp | **3.515%** | **3.679%** | **3.691%** |
148
 
149
- IQ1_M is **8.633 MiB** smaller than IQ2_S and **36.523 MiB** smaller than Q3_K_M. Versus IQ2_S it has a better PPL point estimate and **+0.354 pp** Same top-p, but Mean KLD is **+0.000334 / +2.38%** higher. The differences are smaller than the reported uncertainty scale, so this record treats IQ1_M and IQ2_S as **neighboring low-footprint Pareto points**, not as proof of categorical fidelity superiority.
150
-
151
- ## 9. Release-readiness distinction
152
-
153
- | Evidence layer | IQ2_S | IQ1_M |
154
- |---|---|---|
155
- | Exact tensor structure | PASS | PASS |
156
- | Frozen BF16-paired hard gate | PASS | PASS |
157
- | Engineering reserve gate | PASS | PASS |
158
- | Immutable final SHA256 | PASS | **PASS — local identity frozen** |
159
- | Stock llama.cpp runtime | PASS | **Pending** |
160
- | MTP load/generation/engagement | PASS | **Pending** |
161
- | Remote exact-size + full-SHA + readability | PASS | **Pending** |
162
- | Current public positioning | **Released low-footprint sweet spot** | **Ultra-low-footprint frontier candidate** |
163
-
164
- This distinction is why IQ2_S remains the current released sweet spot: IQ1_M now has quantitative and local-identity PASS, but stock runtime, MTP, and remote verification are not yet closed.
165
-
166
- ## 10. IQ2_S retained release record
167
-
168
- IQ2_S remains fully verified at **18,037,261,056 bytes / 16.798508 GiB**, SHA256 `0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b`, with PPL ratio **1.003457**, Mean KLD **0.014023**, Same top-p **95.516%**, stock runtime PASS, MTP PASS, and remote exact-size/full-SHA/readability PASS.
169
 
170
- ## 11. Artifact identity
171
 
172
- | Artifact | Bytes | GiB | SHA256 | State |
173
- |---|---:|---:|---|---|
174
- | BF16 | **53,808,284,000** | **50.112870** | `11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d` | Verified |
175
- | Q8_0 | **28,595,765,600** | **26.631882** | `6204a47274cfbc0c69c39877fb06615ce842bbab264eea77e2e0a5e3ae2fb8e8` | Verified |
176
- | Q6_K | **22,082,532,096** | **20.565961** | `15d8f856471c4853f6bf0036b2a517426c6cb30a0313cef58a7f9577fd26fe9e` | Verified |
177
- | Q5_K_M | **20,363,666,176** | **18.965142** | `f90ec14d7ec8f084a292413ae0e06483c87ab5e1422ddce51f2e3409d8853b61` | Verified |
178
- | Q4_K_M | **19,650,634,496** | **18.301080** | not frozen in supplied Q4 record | Quantitative candidate |
179
- | Q3_K_M | **18,066,506,496** | **16.825745** | `c2e9539cfb85d87a99605b6914ed602c0fa750164c153ce5574710123aba06fe` | Verified release |
180
- | IQ2_S | **18,037,261,056** | **16.798508** | `0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b` | Verified release |
181
- | **IQ1_M** | **18,028,208,896** | **16.790078** | `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b` | **Quantitative + local identity PASS; runtime/MTP/remote pending** |
182
 
183
- ## 12. Evidence-supported conclusion
184
 
185
- The deployment ladder currently has three confirmed objective-specific sweet spots plus one frontier candidate:
186
 
187
  - **Q6_K — overall quality / efficiency sweet spot.**
188
  - **Q5_K_M — memory-quality sweet spot.**
189
- - **IQ2_S — current released low-footprint sweet spot.**
190
- - **IQ1_M — ultra-low-footprint frontier candidate with frozen local identity; promote to minimum-footprint sweet spot only after stock runtime/MTP/remote identity PASS.**
 
191
 
192
- If IQ1_M reproduces the V2 fidelity record and closes those final release gates, the professionally accurate positioning becomes: **Q6_K overall**, **Q5_K_M memory-quality**, **IQ1_M minimum-footprint**, and **IQ2_S lower-KLD low-footprint alternative**. That is a deployment trade-off classification, not a claim that lower nominal bit count is universally better.
 
7
  **Runtime:** llama.cpp
8
  **Reference precision:** BF16
9
  **Measured ladder:** Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_M, IQ2_S, IQ1_M
10
+ **Parent-model benchmark rerun per quantization tier:** No
11
 
12
+ > This document separates **parent-model capability** from **quantization fidelity**. Parent-model benchmark scores establish the VeriLoop E2 capability record; the measurements below quantify how closely each GGUF tier preserves the canonical BF16 reference under one frozen paired protocol.
13
 
14
  ## 1. Executive decision
15
 
16
+ **Overall sweet spot — Q6_K.** **58.96%** smaller than BF16 with PPL parity within reported uncertainty, Mean KLD **0.004409 ± 0.000953**, and Same top-p **98.204 ± 0.147%**.
17
 
18
+ **Memory-quality sweet spot — Q5_K_M.** **62.16%** smaller than BF16 and **7.78%** smaller than Q6_K, with **+0.4450%** PPL drift, Mean KLD **0.006919 ± 0.000945**, and Same top-p **97.251 ± 0.181%**.
19
 
20
+ **Minimum-footprint sweet spot — IQ1_M.** **16.790078 GiB**, **66.4955% smaller than BF16**, PPL ratio **1.003191 ± 0.002175**, Mean KLD **0.014357 ± 0.001317**, and Same top-p **95.870 ± 0.220%**. Structure, hard fidelity, engineering reserve, stock llama.cpp runtime, and real MTP engagement all pass.
21
 
22
+ **Lower-KLD low-footprint alternative — IQ2_S.** **16.798508 GiB**, Mean KLD **0.014023 ± 0.001408**, and Same top-p **95.516 ± 0.229%**.
 
 
23
 
24
  ## 2. Precision matrix
25
 
26
+ | Tier | Role | Main size | Effective density | Mean PPL | Mean KLD | Same top-p |
27
+ |---|---|---:|---:|---:|---:|---:|
28
+ | BF16 | Reference | **50.112870 GiB** | 16-bit-class | **4.840423 ± 0.119931** | 0 | 100% |
29
+ | Q8_0 | High fidelity | **26.631882 GiB** | **8.50 BPW** | **4.843536 ± 0.120062** | **0.002176 ± 0.000668** | **98.815 ± 0.120%** |
30
+ | **Q6_K** | **Overall sweet spot** | **20.565961 GiB** | **6.57 BPW** | **4.838514 ± 0.119694** | **0.004409 ± 0.000953** | **98.204 ± 0.147%** |
31
+ | **Q5_K_M** | **Memory-quality sweet spot** | **18.965142 GiB** | **6.05 BPW** | **4.861965 ± 0.120630** | **0.006919 ± 0.000945** | **97.251 ± 0.181%** |
32
+ | Q4_K_M | Balanced compact | **18.301080 GiB** | **5.84 BPW** | **4.863760 ± 0.120720** | **0.009700 ± 0.001110** | **96.786 ± 0.195%** |
33
+ | Q3_K_M | Low-footprint alternative | **16.825745 GiB** | **5.37 BPW** | **4.860222 ± 0.120606** | **0.014349 ± 0.001985** | **95.919 ± 0.219%** |
34
+ | IQ2_S | Lower-KLD low-footprint alternative | **16.798508 GiB** | **5.36 BPW** | **4.857159 ± 0.120482** | **0.014023 ± 0.001408** | **95.516 ± 0.229%** |
35
+ | **IQ1_M** | **Minimum-footprint sweet spot** | **16.790078 GiB** | **5.36 BPW** | **4.855870 ± 0.120480** | **0.014357 ± 0.001317** | **95.870 ± 0.220%** |
36
 
37
+ ## 3. Quantization-retention benchmark
38
 
39
  | Tier | PPL ratio | Relative PPL change | Mean KLD | Same top-p | Top-token disagreement | RMS Δp | log-PPL corr. |
40
  |---|---:|---:|---:|---:|---:|---:|---:|
 
46
  | IQ2_S | **1.003457 ± 0.002100** | **+0.3457%** | **0.014023 ± 0.001408** | **95.516 ± 0.229%** | **4.484 pp** | **3.679 ± 0.231%** | **99.64%** |
47
  | **IQ1_M** | **1.003191 ± 0.002175** | **+0.3191%** | **0.014357 ± 0.001317** | **95.870 ± 0.220%** | **4.130 pp** | **3.691 ± 0.205%** | **99.62%** |
48
 
49
+ **Claim boundary.** These are quantization-retention measurements, not downstream task-score loss. The nine parent-model benchmarks were not independently rerun for every quantization tier.
50
 
51
  ## 4. Frozen BF16-paired protocol
52
 
 
63
  | Batch / micro-batch | 512 / 512 |
64
  | Evaluator | `llama-perplexity` |
65
  | Reference method | `--kl-divergence-base` |
66
+ | Quantized comparison method | `--kl-divergence` |
67
  | BF16 logits reused | Yes |
68
  | llama.cpp revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` |
69
 
 
79
  | Mean Δp | 0 | **−0.042 ± 0.041%** | small directional bias |
80
  | log-PPL correlation | 100% | **99.62%** | high agreement |
81
 
82
+ ### KLD distribution
83
 
84
  | Statistic | Value |
85
  |---|---:|
 
93
 
94
  ## 6. IQ1_M exact tensor policy
95
 
96
+ IQ1_M differs from IQ2_S by exactly one tensor:
97
 
98
  - `blk.1.ffn_down.weight`: **IQ2_S → IQ1_M**
99
+ - `blk.0.ffn_down.weight`: **IQ2_S**
100
+ - `blk.3.ffn_down.weight`: **IQ2_S**
101
 
102
  | Precision | Count |
103
  |---|---:|
 
109
  | Q6_K | **2** |
110
  | **Total** | **851** |
111
 
112
+ | Build item | Value |
113
  |---|---|
114
  | Source | Canonical BF16 |
115
  | Low-bit requantization | **No** |
 
121
  | End-to-end quantization timer | **168.26090 s** |
122
  | SHA256 | `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b` |
123
 
124
+ ## 7. IQ1_M final validation
 
 
 
 
 
 
 
125
 
126
+ | Validation | Requirement | Measured | Result |
127
+ |---|---|---|---|
128
+ | PPL ratio hard / reserve | ≤1.0150 / ≤1.0135 | **1.003191** | **PASS / PASS** |
129
+ | Mean KLD hard / reserve | ≤0.0150 / ≤0.0145 | **0.014357** | **PASS / PASS** |
130
+ | Same top-p hard / reserve | ≥95.0% / ≥95.5% | **95.870%** | **PASS / PASS** |
131
+ | Tensor structure | exact policy | **851; 0 mismatch** | **PASS** |
132
+ | Stock llama.cpp runtime | real HTTP generation | **200; non-empty output** | **PASS** |
133
+ | MTP runtime | real HTTP generation | **200; non-empty output** | **PASS** |
134
+ | MTP engagement | generated > 0; accepted > 0 | **76 / 104 accepted (73.0769%)** | **PASS** |
135
+ | Acceptance-rate consistency | exact record | **0.73076923** | **PASS** |
136
+ | Main/MTP output audit | advisory | **IDENTICAL** | **PASS** |
137
 
138
+ ## 8. Low-footprint comparison
 
 
139
 
140
  | Metric | Q3_K_M | IQ2_S | IQ1_M |
141
  |---|---:|---:|---:|
 
145
  | Same top-p | **95.919%** | **95.516%** | **95.870%** |
146
  | RMS Δp | **3.515%** | **3.679%** | **3.691%** |
147
 
148
+ IQ1_M is **8.633 MiB** smaller than IQ2_S and **36.523 MiB** smaller than Q3_K_M. IQ2_S has the lowest Mean KLD, while IQ1_M has the smallest footprint, the lowest PPL ratio, and higher Same top-p than IQ2_S. The differences remain small relative to the reported uncertainty scale.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
149
 
150
+ ## 9. Frozen artifact identities
151
 
152
+ | Artifact | Bytes | GiB | SHA256 |
153
+ |---|---:|---:|---|
154
+ | BF16 | **53,808,284,000** | **50.112870** | `11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d` |
155
+ | Q8_0 | **28,595,765,600** | **26.631882** | `6204a47274cfbc0c69c39877fb06615ce842bbab264eea77e2e0a5e3ae2fb8e8` |
156
+ | Q6_K | **22,082,532,096** | **20.565961** | `15d8f856471c4853f6bf0036b2a517426c6cb30a0313cef58a7f9577fd26fe9e` |
157
+ | Q5_K_M | **20,363,666,176** | **18.965142** | `f90ec14d7ec8f084a292413ae0e06483c87ab5e1422ddce51f2e3409d8853b61` |
158
+ | Q3_K_M | **18,066,506,496** | **16.825745** | `c2e9539cfb85d87a99605b6914ed602c0fa750164c153ce5574710123aba06fe` |
159
+ | IQ2_S | **18,037,261,056** | **16.798508** | `0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b` |
160
+ | **IQ1_M** | **18,028,208,896** | **16.790078** | `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b` |
 
161
 
162
+ Q4_K_M remains part of the measured retention ladder; this identity table intentionally lists only artifacts with a frozen SHA256 in the supplied release record.
163
 
164
+ ## 10. Conclusion
165
 
166
  - **Q6_K — overall quality / efficiency sweet spot.**
167
  - **Q5_K_M — memory-quality sweet spot.**
168
+ - **IQ1_M — minimum-footprint sweet spot.**
169
+ - **IQ2_S — lower-KLD low-footprint alternative.**
170
+ - **Q3_K_M — adjacent low-footprint alternative.**
171
 
172
+ This classification is deployment-oriented. It does not imply that lower nominal bit count is universally superior.
README.md CHANGED
@@ -39,7 +39,7 @@ tags:
39
 
40
  <p align="center">
41
  <strong>Official llama.cpp distribution of VeriLoop E2</strong><br>
42
- <em>BF16 reference · Q8_0 high fidelity · Q6_K overall sweet spot · Q5_K_M memory-quality sweet spot · IQ2_S released low-footprint sweet spot · IQ1_M ultra-low-footprint frontier</em><br><br>
43
  <strong>27B post-trained model for code, mathematics, and physics · 262K native context · Apache License 2.0</strong><br>
44
  <strong>Developed by Tsinghua SIGS Robot Lab · Libo Wang</strong>
45
  </p>
@@ -49,8 +49,7 @@ tags:
49
  <img src="https://img.shields.io/badge/Format-GGUF-111827?style=flat-square" alt="Format: GGUF">
50
  <img src="https://img.shields.io/badge/Q6__K-Overall%20Sweet%20Spot-0A7F6F?style=flat-square" alt="Q6_K: Overall Sweet Spot">
51
  <img src="https://img.shields.io/badge/Q5__K__M-Memory--Quality%20Sweet%20Spot-0F766E?style=flat-square" alt="Q5_K_M: Memory Quality Sweet Spot">
52
- <img src="https://img.shields.io/badge/IQ2__S-Low--Footprint%20Sweet%20Spot-475569?style=flat-square" alt="IQ2_S: Low-Footprint Sweet Spot">
53
- <img src="https://img.shields.io/badge/IQ1__M-Ultra--Low--Footprint%20Frontier-334155?style=flat-square" alt="IQ1_M: Ultra-Low-Footprint Frontier">
54
  <img src="https://img.shields.io/badge/Runtime-llama.cpp-0A7F6F?style=flat-square" alt="Runtime: llama.cpp">
55
  </p>
56
 
@@ -63,48 +62,45 @@ tags:
63
 
64
  ---
65
 
66
- ## Which file should I download?
67
 
68
- | Priority | Recommended file | Main size | Measured retention | Use when |
69
- |---|---|---:|---|---|
70
- | **Default** | **`VeriLoop-E2-Q6_K.gguf`** | **20.566 GiB** | BF16-paired PPL parity within uncertainty; KLD **0.004409**; Same top-p **98.204%** | Strongest overall quality / footprint balance |
71
- | **Tighter memory** | **`VeriLoop-E2-Q5_K_M.gguf`** | **18.965 GiB** | PPL **+0.4450%**; KLD **0.006919**; Same top-p **97.251%** | Smaller deployment with a wider fidelity margin |
72
- | **Released low-footprint sweet spot** | **`VeriLoop-E2-IQ2_S.gguf`** | **16.799 GiB** | PPL **+0.3457%**; KLD **0.014023**; Same top-p **95.516%** | Smallest fully release-verified low-footprint tier |
73
- | **Ultra-low-footprint frontier candidate** | **`VeriLoop-E2-IQ1_M.gguf`** | **16.790 GiB** | PPL **+0.3191%**; KLD **0.014357**; Same top-p **95.870%** | Smallest measured candidate; quantitative fidelity and local identity verified |
74
- | **Low-footprint alternative** | `VeriLoop-E2-Q3_K_M.gguf` | **16.826 GiB** | PPL **+0.4090%**; KLD **0.014349**; Same top-p **95.919%** | Adjacent fully validated option |
75
- | **High fidelity** | `VeriLoop-E2-Q8_0.gguf` | **26.632 GiB** | PPL **+0.0643%**; KLD **0.002176**; Same top-p **98.815%** | Distributional fidelity matters more than footprint |
76
- | **Reference** | `VeriLoop-E2-BF16.gguf` | **50.113 GiB** | Canonical reference | Reproducibility, attribution, quantization research |
 
77
 
78
- ### Current sweet spots and frontier
79
 
80
- **Q6_K remains the overall quality / efficiency sweet spot.** It is **58.96% smaller than BF16** while its PPL point estimate remains statistically consistent with BF16 parity under the frozen paired protocol.
81
 
82
- **Q5_K_M remains the memory-quality sweet spot.** It offers a materially smaller footprint than Q6_K while keeping substantially wider KLD and Same-top margins than the sub-17 GiB frontier.
83
 
84
- **IQ2_S remains the current released low-footprint sweet spot.** It is fully verified through structure, BF16-paired fidelity, engineering reserve, stock llama.cpp runtime, MTP engagement, immutable SHA256, and remote readability.
85
 
86
- **IQ1_M is the current ultra-low-footprint frontier candidate.** Its measured V2 artifact is **16.790078 GiB**, **66.4955% smaller than BF16**, **8.633 MiB / 0.0502% smaller than IQ2_S**, and **36.523 MiB / 0.2120% smaller than Q3_K_M**. Against IQ2_S, IQ1_M has a slightly better PPL ratio (**1.003191 vs 1.003457**) and higher Same top-p (**95.870% vs 95.516%**), but a slightly higher Mean KLD (**0.014357 vs 0.014023**). The uncertainty bands substantially overlap, so this release does **not** claim that either low-footprint tier is universally superior in fidelity.
87
-
88
- **Positioning.** IQ1_M defines the current minimum-footprint frontier, while IQ2_S remains the released low-footprint balance point / lower-KLD alternative.
89
-
90
- > **Naming note:** `VeriLoop-E2-IQ1_M.gguf` is not a uniform 1-bit model. Its measured policy is **353 F32 + 1 IQ1_M + 2 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K = 851 tensors**, with **5.36 effective BPW**. `VeriLoop-E2-IQ2_S.gguf` is likewise a mixed-precision release rather than a uniform 2-bit model.
91
 
92
  ## Precision ladder
93
 
94
- | Tier | Role | Main size | Reduction vs BF16 | Effective density | Release state |
95
- |---|---|---:|---:|---:|---|
96
- | **BF16** | Canonical reference | **53.808 GB / 50.113 GiB** | — | 16-bit-class | Verified reference |
97
- | **Q8_0** | High-fidelity local | **28.596 GB / 26.632 GiB** | **46.86%** | **8.50 BPW** | Verified |
98
- | **Q6_K** | **Overall sweet spot** | **22.083 GB / 20.566 GiB** | **58.96%** | **6.57 BPW** | Verified / Release Quality PASS |
99
- | **Q5_K_M** | **Memory-quality sweet spot** | **20.364 GB / 18.965 GiB** | **62.16%** | **6.05 BPW** | Verified / Release Quality PASS |
100
- | **Q4_K_M** | Intermediate mixed-precision candidate | **19.651 GB / 18.301 GiB** | **63.48%** | **5.84 BPW** | Quantitatively validated candidate |
101
- | **Q3_K_M** | Adjacent low-footprint alternative | **18.067 GB / 16.826 GiB** | **66.42%** | **5.37 BPW** | Full release PASS |
102
- | **IQ2_S** | **Released low-footprint sweet spot** | **18.037 GB / 16.799 GiB** | **66.48%** | **5.36 BPW** | **Full release PASS** |
103
- | **IQ1_M** | **Ultra-low-footprint frontier candidate** | **18.028 GB / 16.790 GiB** | **66.50%** | **5.36 BPW** | **Quantitative fidelity PASS; local identity verified** |
104
 
105
  ## Quantization-retention benchmark
106
 
107
- All tiers below use the same frozen BF16 logits and the same paired protocol.
108
 
109
  | Tier | Mean PPL | PPL ratio vs BF16 | Relative PPL change | Mean KLD | Same top-p | log-PPL correlation |
110
  |---|---:|---:|---:|---:|---:|---:|
@@ -117,15 +113,11 @@ All tiers below use the same frozen BF16 logits and the same paired protocol.
117
  | **IQ2_S** | **4.857159 ± 0.120482** | **1.003457 ± 0.002100** | **+0.3457%** | **0.014023 ± 0.001408** | **95.516 ± 0.229%** | **99.64%** |
118
  | **IQ1_M** | **4.855870 ± 0.120480** | **1.003191 ± 0.002175** | **+0.3191%** | **0.014357 ± 0.001317** | **95.870 ± 0.220%** | **99.62%** |
119
 
120
- ### What “performance loss” means here
121
-
122
- The **+0.3191% IQ1_M** and **+0.3457% IQ2_S** figures are **PPL drift under the frozen BF16-paired quantization benchmark**. They are not claims that SWE-bench, Terminal-Bench, DeepSWE, AIME, GPQA, Apex, coding-agent capability, mathematics, or physics scores fall by those percentages. The parent nine-benchmark campaign has **not** been rerun independently for each quant tier, so no quant-specific downstream score-loss percentage is claimed.
123
 
124
- For the low-footprint frontier, PPL must be read together with KLD, Same top-p, RMS Δp, and uncertainty. IQ1_M improves the PPL point estimate and Same top-p versus IQ2_S, while IQ2_S has **2.38% lower Mean KLD** and slightly lower RMS Δp. These differences are small relative to their reported uncertainties; the professional interpretation is **trade-off parity**, not a claim of categorical quality superiority.
125
 
126
- ## IQ1_M — ultra-low-footprint frontier candidate
127
-
128
- The measured IQ1_M policy is a **single-factor step from the successful IQ2_S release**: only `blk.1.ffn_down.weight` changes from IQ2_S to IQ1_M. Every other tensor assignment remains on the Q2 policy.
129
 
130
  | Precision | Tensor assignment | Count |
131
  |---|---|---:|
@@ -137,112 +129,44 @@ The measured IQ1_M policy is a **single-factor step from the successful IQ2_S re
137
  | IQ1_M | `blk.1.ffn_down.weight` | **1** |
138
  | **Total** | | **851** |
139
 
140
- ### IQ1_M measured retention
141
-
142
- | Metric | BF16 | IQ1_M | Observed drift |
143
- |---|---:|---:|---:|
144
- | Mean PPL | **4.840423 ± 0.119931** | **4.855870 ± 0.120480** | **+0.015446 ± 0.010522** |
145
- | PPL ratio | 1.000000 | **1.003191 ± 0.002175** | **+0.3191%** |
146
- | Mean KLD | 0 | **0.014357 ± 0.001317** | lower is better |
147
- | Same top-p | 100% | **95.870 ± 0.220%** | **4.130 pp disagreement** |
148
- | RMS Δp | 0 | **3.691 ± 0.205%** | distributional-noise scale |
149
- | Mean Δp | 0 | **−0.042 ± 0.041%** | small directional bias |
150
- | log-PPL correlation | 100% | **99.62%** | high agreement |
151
-
152
- ### IQ1_M final quantitative gates
153
-
154
- | Gate | Hard threshold | Reserve threshold | Measured V2 | Quantitative result |
155
- |---|---:|---:|---:|---|
156
- | PPL ratio | ≤ **1.0150** | ≤ **1.0135** | **1.003191** | **PASS / PASS** |
157
- | Mean KLD | ≤ **0.0150** | ≤ **0.0145** | **0.014357** | **PASS / PASS** |
158
- | Same top-p | ≥ **95.0%** | ≥ **95.5%** | **95.870%** | **PASS / PASS** |
159
- | Exact tensor structure | declared policy | declared policy | **851 tensors; 0 mismatch** | **PASS** |
160
-
161
- The original V2 script rejected this artifact only because an additional experimental promotion rule demanded that **every Q1 point estimate also dominate IQ2_S**, including Mean KLD. That requirement has been removed from V3; the underlying hard and engineering-reserve thresholds are unchanged. The V2 artifact itself was deleted after that rejection, so its SHA256 is not represented as a frozen release identity. The accepted rebuild has frozen the identical tensor policy and local SHA256; runtime/MTP/remote identity must still close before public release.
162
-
163
- ### IQ1_M vs IQ2_S — low-footprint frontier
164
-
165
- | Metric | IQ2_S | IQ1_M | IQ1_M delta |
166
- |---|---:|---:|---:|
167
- | Main size | **16.798508 GiB** | **16.790078 GiB** | **−8.633 MiB / −0.0502%** |
168
- | PPL ratio | **1.003457** | **1.003191** | **−0.000266** |
169
- | Mean KLD | **0.014023** | **0.014357** | **+0.000334 / +2.38%** |
170
- | Same top-p | **95.516%** | **95.870%** | **+0.354 pp** |
171
- | RMS Δp | **3.679%** | **3.691%** | **+0.012 pp** |
172
- | log-PPL correlation | **99.64%** | **99.62%** | **−0.02 pp** |
173
-
174
- **Interpretation:** IQ1_M is not merely “worse Q2.” It is a neighboring Pareto point: **smaller, slightly better on PPL and top-token agreement, slightly worse on KLD/RMS**, with overlapping uncertainty. Its major deployment advantage is the **absolute 16.790 GiB footprint / 66.4955% reduction from BF16**; the incremental saving over IQ2_S itself is only **8.633 MiB**, so that incremental difference should not be exaggerated.
175
-
176
- ## IQ2_S — root-cause-guided low-footprint sweet spot
177
-
178
- IQ2_S is derived **directly from the canonical BF16 GGUF**. It preserves the successful Q3 precision spine and changes only the three root-cause-selected FFN-down tensors from Q3_K to IQ2_S.
179
-
180
- | Precision | Tensor assignment | Count |
181
- |---|---|---:|
182
- | F32 | Non-quantized tensors retained by GGUF conversion | **353** |
183
- | Q6_K | `output.weight`, `token_embd.weight` | **2** |
184
- | Q5_K | Remaining quantized internal tensors | **429** |
185
- | Q4_K | all `*.ffn_up.weight` | **64** |
186
- | IQ2_S | `blk.0.ffn_down.weight`, `blk.1.ffn_down.weight`, `blk.3.ffn_down.weight` | **3** |
187
- | **Total** | | **851** |
188
-
189
- ### Q2 development path
190
-
191
- | Stage | Targeted change | Main size | PPL ratio | Mean KLD | Same top-p | Decision |
192
- |---|---|---:|---:|---:|---:|---|
193
- | IQ2_S V1 | Four low-sensitivity `ffn_up` tensors → IQ2_S | **16.809533 GiB** | **1.008995** | **0.014951** | **95.809%** | **FAIL reserve KLD** |
194
- | **IQ2_S V2 final** | **Q3 spine; three selected `ffn_down` tensors → IQ2_S** | **16.798508 GiB** | **1.003457** | **0.014023** | **95.516%** | **PASS hard + reserve** |
195
-
196
- V1 was rejected rather than promoted: it saved only **0.0964%** versus Q3 and exceeded the frozen **0.0145** Mean-KLD reserve. V2 changes the tensor family instead of relaxing the threshold.
197
 
198
- ### IQ2_S final gates
 
 
 
 
 
 
 
 
 
199
 
200
- | Gate | Hard threshold | Reserve threshold | Measured | Result |
201
- |---|---:|---:|---:|---|
202
- | PPL ratio | ≤ **1.0150** | ≤ **1.0135** | **1.003457** | **PASS / PASS** |
203
- | Mean KLD | ≤ **0.0150** | ≤ **0.0145** | **0.014023** | **PASS / PASS** |
204
- | Same top-p | ≥ **95.0%** | ≥ **95.5%** | **95.516%** | **PASS / PASS** |
205
- | Exact tensor structure | declared policy | declared policy | **851 tensors; 0 mismatch** | **PASS** |
206
- | Stock llama.cpp main runtime | required | required | **PASS** | **PASS** |
207
- | MTP load + generation | required | required | **PASS** | **PASS** |
208
- | Real MTP engagement | >0 accepted | >0 accepted | **116 / 331 accepted (35.045%)** | **PASS** |
209
-
210
- Reserve headroom is **0.010043** on PPL ratio, **0.000477** on Mean KLD, and **0.016 pp** on Same top-p. The final value is intentionally exposed: **this is a validated frontier, not evidence that further low-bit expansion is safe.**
211
-
212
- ## IQ2_S immutable identity
213
 
214
  | Property | Value |
215
  |---|---|
216
- | Filename | `VeriLoop-E2-IQ2_S.gguf` |
217
- | Exact bytes | **18,037,261,056** |
218
- | Binary size | **16.798508 GiB** |
219
- | SHA256 | `0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b` |
220
  | Effective density | **5.36 BPW** |
221
  | Tensor count | **851** |
222
- | Tensor distribution | **353 F32 + 3 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K** |
223
- | Quantized size reported by quantizer | **17,191.19 MiB** |
224
- | Quantization time | **168.08618 s** |
225
- | llama.cpp QC/runtime revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` |
226
- | Hugging Face remote identity | **PASS — exact size + full SHA256 + remote readability** |
227
-
228
- ## Q3_K_M — adjacent low-footprint alternative
229
 
230
- Q3 remains a valid release. It has **0.403 pp higher Same top-p** than IQ2_S, while IQ2_S is **27.891 MiB / 0.162% smaller**, with a lower PPL ratio and **2.27% lower Mean KLD**. For minimum footprint, IQ2_S is the preferred frontier point; Q3 remains the adjacent option when top-token agreement is prioritized.
231
 
232
- ## Optional MTP draft
233
-
234
- IQ2_S and Q3 were validated with the same 18-tensor Q4_0 MTP object:
235
-
236
- | Property | Value |
237
- |---|---|
238
- | MTP bytes | **2,008,056,096 / 1.870 GiB** |
239
- | MTP SHA256 | `0e7f2dfe254f3a5d195d105832411acc1007fd7bb9ac980fe9aae4a1f290d671` |
240
- | IQ2_S draft tokens generated | **331** |
241
- | IQ2_S draft tokens accepted | **116** |
242
- | IQ2_S acceptance rate | **0.35045 / 35.045%** |
243
- | Mean draft length | **3.76** |
244
 
245
- A duplicate Q2-specific MTP file is not required for this validated path. Main-only and main+MTP continuations diverged in the deterministic audit; that result is retained as an advisory. MTP engagement is established by explicit draft-model loading plus non-zero generated and accepted draft tokens.
246
 
247
  ## Frozen BF16-paired protocol
248
 
@@ -260,13 +184,13 @@ A duplicate Q2-specific MTP file is not required for this validated path. Main-o
260
  | Batch / micro-batch | **512 / 512** |
261
  | Evaluator | `llama-perplexity` |
262
  | Reference logits | `--kl-divergence-base` |
263
- | Candidate comparison | `--kl-divergence` |
264
  | BF16 logits reused across tiers | **Yes** |
265
  | llama.cpp revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` |
266
 
267
  ## Parent-model benchmark record
268
 
269
- These scores identify the **parent VeriLoop E2 release**. They provide capability context and are **not relabeled as quant-specific re-runs**.
270
 
271
  | Benchmark | VeriLoop E2 parent score | Public evidence |
272
  |---|---:|---|
@@ -286,66 +210,41 @@ These scores identify the **parent VeriLoop E2 release**. They provide capabilit
286
  # Overall sweet spot
287
  hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-Q6_K.gguf --local-dir .
288
 
289
- # Released low-footprint sweet spot
290
- hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-IQ2_S.gguf --local-dir .
291
- ```
292
 
293
- > **IQ1_M is staged as the ultra-low-footprint frontier artifact.** Its quantitative evidence is PASS and its local SHA256 is frozen as `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`.
294
-
295
- Run IQ2_S:
296
-
297
- ```bash
298
- llama-server -m ./VeriLoop-E2-IQ2_S.gguf -ngl 99 -c 32768 --host 127.0.0.1 --port 8080
 
299
  ```
300
 
301
- IQ2_S + validated MTP companion:
302
 
303
- ```bash
304
- hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF mtp-VeriLoop-E2-Q5_K_M.gguf --local-dir .
305
-
306
- llama-server -m ./VeriLoop-E2-IQ2_S.gguf --model-draft ./mtp-VeriLoop-E2-Q5_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 8 --gpu-layers-draft 99 -ngl 99 -c 2048 --host 127.0.0.1 --port 8080
307
- ```
308
-
309
- ## Memory and throughput guidance
310
-
311
- | Variant | Main file | Position |
312
  |---|---:|---|
313
  | BF16 | **50.113 GiB** | Reference |
314
  | Q8_0 | **26.632 GiB** | High fidelity |
315
  | **Q6_K** | **20.566 GiB** | **Overall sweet spot** |
316
  | **Q5_K_M** | **18.965 GiB** | **Memory-quality sweet spot** |
317
- | Q4_K_M | **18.301 GiB** | Intermediate quantitative candidate |
318
- | Q3_K_M | **16.826 GiB** | Adjacent low-footprint release |
319
- | **IQ2_S** | **16.799 GiB** | **Released low-footprint sweet spot** |
320
- | **IQ1_M** | **16.790 GiB** | **Ultra-low-footprint frontier candidate** |
321
-
322
- A **16.79 GiB model file does not imply full offload on a 16 GiB GPU**. Runtime memory also includes KV cache, compute buffers, allocator overhead, and optional MTP weights.
323
 
324
- A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128 41.999677 tok/s**), but there is no frozen paired BF16 throughput campaign. No universal speedup percentage is claimed for Q6/Q5/Q4/Q3/IQ2_S/IQ1_M.
325
 
326
- ## Release positioning
327
-
328
- | Tier | Positioning |
329
- |---|---|
330
- | BF16 | Canonical reference |
331
- | Q8_0 | High-fidelity local |
332
- | **Q6_K** | **Overall quality / efficiency sweet spot** |
333
- | **Q5_K_M** | **Memory-quality sweet spot** |
334
- | Q4_K_M | Intermediate mixed-precision candidate |
335
- | Q3_K_M | Adjacent low-footprint release |
336
- | **IQ2_S** | **Current released low-footprint sweet spot / lower-KLD frontier point** |
337
- | **IQ1_M** | **Ultra-low-footprint frontier candidate** |
338
 
339
  ## Measurement boundaries
340
 
341
- - PPL/KLD/token metrics are **quantization-retention measurements**, not universal percentages of downstream capability loss.
342
- - IQ1_M's **+0.3191% PPL** does not mean a 0.3191% drop on SWE-bench, Terminal-Bench, AIME, GPQA, or another downstream benchmark.
343
- - IQ2_S's **+0.3457% PPL** likewise does not imply a 0.3457% downstream-score loss.
344
- - IQ1_M and IQ2_S are both mixed-precision artifacts; the nominal tier name is not the model-wide effective bit width.
345
- - IQ1_M has a frozen local SHA256 (`e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`) and quantitative PASS evidence.
346
- - The IQ1_M vs IQ2_S point-estimate differences are small relative to the reported uncertainty; this repository does not claim categorical fidelity superiority for either frontier point.
347
  - Native 262K context comes from the parent configuration; practical context depends on runtime memory.
348
- - MTP acceptance is prompt/workload-dependent.
349
  - The model can still produce incorrect code, mathematics, scientific reasoning, or commands.
350
 
351
  ## Source model and evidence
@@ -355,7 +254,7 @@ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128
355
  | Parent model | [VeriLoop E2](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2) |
356
  | GGUF repository | [VeriLoop E2 GGUF](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF) |
357
  | Quantization quality record | [Quantization Quality](./QUANTIZATION_QUALITY.md) |
358
- | Artifact identity / release state | [Release Manifest](./RELEASE_MANIFEST.json) |
359
  | Technical report | [OpenReview](https://openreview.net/forum?id=P6FIQILHwX&noteId=P6FIQILHwX) |
360
  | Evaluation evidence | [VeriLoop E2 Evaluation Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence) |
361
  | Riemann ζ artifact | [Public artifact](https://github.com/brucewang123456789/GeniusTrail/tree/VeriLoop-E2/riemann-hypothesis) |
@@ -378,4 +277,4 @@ The VeriLoop E2 model weights and this GGUF distribution are released under the
378
  }
379
  ```
380
 
381
- For quantization-specific comparisons, identify the exact GGUF filename and corresponding release-manifest identity.
 
39
 
40
  <p align="center">
41
  <strong>Official llama.cpp distribution of VeriLoop E2</strong><br>
42
+ <em>BF16 reference · Q8_0 high fidelity · Q6_K overall sweet spot · Q5_K_M memory-quality sweet spot · IQ1_M minimum-footprint sweet spot</em><br><br>
43
  <strong>27B post-trained model for code, mathematics, and physics · 262K native context · Apache License 2.0</strong><br>
44
  <strong>Developed by Tsinghua SIGS Robot Lab · Libo Wang</strong>
45
  </p>
 
49
  <img src="https://img.shields.io/badge/Format-GGUF-111827?style=flat-square" alt="Format: GGUF">
50
  <img src="https://img.shields.io/badge/Q6__K-Overall%20Sweet%20Spot-0A7F6F?style=flat-square" alt="Q6_K: Overall Sweet Spot">
51
  <img src="https://img.shields.io/badge/Q5__K__M-Memory--Quality%20Sweet%20Spot-0F766E?style=flat-square" alt="Q5_K_M: Memory Quality Sweet Spot">
52
+ <img src="https://img.shields.io/badge/IQ1__M-Minimum%20Footprint-334155?style=flat-square" alt="IQ1_M: Minimum Footprint Sweet Spot">
 
53
  <img src="https://img.shields.io/badge/Runtime-llama.cpp-0A7F6F?style=flat-square" alt="Runtime: llama.cpp">
54
  </p>
55
 
 
62
 
63
  ---
64
 
65
+ ## Model variants
66
 
67
+ | Use case | File | Main size | BF16-paired retention |
68
+ |---|---|---:|---|
69
+ | **Default / overall balance** | **`VeriLoop-E2-Q6_K.gguf`** | **20.566 GiB** | PPL parity within uncertainty; KLD **0.004409**; Same top-p **98.204%** |
70
+ | **Memory-quality balance** | **`VeriLoop-E2-Q5_K_M.gguf`** | **18.965 GiB** | PPL **+0.4450%**; KLD **0.006919**; Same top-p **97.251%** |
71
+ | **Minimum footprint** | **`VeriLoop-E2-IQ1_M.gguf`** | **16.790 GiB** | PPL **+0.3191%**; KLD **0.014357**; Same top-p **95.870%** |
72
+ | Lower-KLD low-footprint alternative | `VeriLoop-E2-IQ2_S.gguf` | **16.799 GiB** | PPL **+0.3457%**; KLD **0.014023**; Same top-p **95.516%** |
73
+ | Low-footprint alternative | `VeriLoop-E2-Q3_K_M.gguf` | **16.826 GiB** | PPL **+0.4090%**; KLD **0.014349**; Same top-p **95.919%** |
74
+ | Balanced compact | `VeriLoop-E2-Q4_K_M.gguf` | **18.301 GiB** | PPL **+0.4821%**; KLD **0.009700**; Same top-p **96.786%** |
75
+ | High fidelity | `VeriLoop-E2-Q8_0.gguf` | **26.632 GiB** | PPL **+0.0643%**; KLD **0.002176**; Same top-p **98.815%** |
76
+ | Reference | `VeriLoop-E2-BF16.gguf` | **50.113 GiB** | Canonical BF16 reference |
77
 
78
+ ### Recommended deployment points
79
 
80
+ **Q6_K — overall quality / efficiency sweet spot.** It is **58.96% smaller than BF16** while remaining statistically consistent with BF16 PPL parity under the frozen paired protocol.
81
 
82
+ **Q5_K_M — memory-quality sweet spot.** It reduces the main-file footprint to **18.965 GiB** while preserving wider KLD and Same-top margins than the sub-17 GiB variants.
83
 
84
+ **IQ1_M — minimum-footprint sweet spot.** It is **16.790078 GiB**, **66.4955% smaller than BF16**, and passes the frozen hard gate, engineering-reserve gate, stock llama.cpp runtime validation, and real MTP engagement validation. IQ2_S remains the lower-KLD low-footprint alternative.
85
 
86
+ > **Naming note:** `VeriLoop-E2-IQ1_M.gguf` is a mixed-precision artifact, not a uniform 1-bit model. Its measured tensor policy is **353 F32 + 1 IQ1_M + 2 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K = 851 tensors**, with **5.36 effective BPW**. `VeriLoop-E2-IQ2_S.gguf` is likewise mixed precision rather than uniform 2-bit quantization.
 
 
 
 
87
 
88
  ## Precision ladder
89
 
90
+ | Tier | Role | Main size | Reduction vs BF16 | Effective density |
91
+ |---|---|---:|---:|---:|
92
+ | **BF16** | Canonical reference | **53.808 GB / 50.113 GiB** | — | 16-bit-class |
93
+ | **Q8_0** | High fidelity | **28.596 GB / 26.632 GiB** | **46.86%** | **8.50 BPW** |
94
+ | **Q6_K** | **Overall sweet spot** | **22.083 GB / 20.566 GiB** | **58.96%** | **6.57 BPW** |
95
+ | **Q5_K_M** | **Memory-quality sweet spot** | **20.364 GB / 18.965 GiB** | **62.16%** | **6.05 BPW** |
96
+ | **Q4_K_M** | Balanced compact | **19.651 GB / 18.301 GiB** | **63.48%** | **5.84 BPW** |
97
+ | **Q3_K_M** | Low-footprint alternative | **18.067 GB / 16.826 GiB** | **66.42%** | **5.37 BPW** |
98
+ | **IQ2_S** | Lower-KLD low-footprint alternative | **18.037 GB / 16.799 GiB** | **66.48%** | **5.36 BPW** |
99
+ | **IQ1_M** | **Minimum-footprint sweet spot** | **18.028 GB / 16.790 GiB** | **66.50%** | **5.36 BPW** |
100
 
101
  ## Quantization-retention benchmark
102
 
103
+ All measured tiers use the same frozen BF16 logits and the same paired protocol.
104
 
105
  | Tier | Mean PPL | PPL ratio vs BF16 | Relative PPL change | Mean KLD | Same top-p | log-PPL correlation |
106
  |---|---:|---:|---:|---:|---:|---:|
 
113
  | **IQ2_S** | **4.857159 ± 0.120482** | **1.003457 ± 0.002100** | **+0.3457%** | **0.014023 ± 0.001408** | **95.516 ± 0.229%** | **99.64%** |
114
  | **IQ1_M** | **4.855870 ± 0.120480** | **1.003191 ± 0.002175** | **+0.3191%** | **0.014357 ± 0.001317** | **95.870 ± 0.220%** | **99.62%** |
115
 
116
+ These figures measure **quantization retention against the BF16 reference**. They are not downstream benchmark-score loss percentages. The nine parent-model benchmarks were not independently rerun for every quantization tier.
 
 
117
 
118
+ ## IQ1_M validation record
119
 
120
+ IQ1_M changes exactly one tensor relative to the IQ2_S precision policy: `blk.1.ffn_down.weight` moves from IQ2_S to IQ1_M. All other tensor assignments remain unchanged.
 
 
121
 
122
  | Precision | Tensor assignment | Count |
123
  |---|---|---:|
 
129
  | IQ1_M | `blk.1.ffn_down.weight` | **1** |
130
  | **Total** | | **851** |
131
 
132
+ ### Fidelity and runtime validation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
133
 
134
+ | Validation | Requirement | Measured | Result |
135
+ |---|---|---|---|
136
+ | PPL ratio hard / reserve | ≤ **1.0150** / ≤ **1.0135** | **1.003191** | **PASS / PASS** |
137
+ | Mean KLD hard / reserve | ≤ **0.0150** / ≤ **0.0145** | **0.014357** | **PASS / PASS** |
138
+ | Same top-p hard / reserve | ≥ **95.0%** / ≥ **95.5%** | **95.870%** | **PASS / PASS** |
139
+ | Exact tensor structure | 851 tensors; declared policy | **851 tensors; 0 mismatch** | **PASS** |
140
+ | Stock llama.cpp main runtime | Successful real generation | HTTP **200**, non-empty generation | **PASS** |
141
+ | MTP runtime | Successful real generation | HTTP **200**, non-empty generation | **PASS** |
142
+ | Real MTP engagement | generated > 0; accepted > 0 | **76 accepted / 104 generated (73.0769%)** | **PASS** |
143
+ | Main vs MTP deterministic audit | Advisory | Identical output SHA256 | **IDENTICAL** |
144
 
145
+ **IQ1_M artifact identity**
 
 
 
 
 
 
 
 
 
 
 
 
146
 
147
  | Property | Value |
148
  |---|---|
149
+ | Filename | `VeriLoop-E2-IQ1_M.gguf` |
150
+ | Exact bytes | **18,028,208,896** |
151
+ | Binary size | **16.790078 GiB** |
152
+ | SHA256 | `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b` |
153
  | Effective density | **5.36 BPW** |
154
  | Tensor count | **851** |
155
+ | Quantizer-reported size | **17,182.55 MiB** |
156
+ | Quantizer time | **167.77645 s** |
157
+ | llama.cpp validation revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` |
 
 
 
 
158
 
159
+ ### Low-footprint comparison
160
 
161
+ | Metric | Q3_K_M | IQ2_S | IQ1_M |
162
+ |---|---:|---:|---:|
163
+ | Main size | **16.825745 GiB** | **16.798508 GiB** | **16.790078 GiB** |
164
+ | PPL ratio | **1.004090** | **1.003457** | **1.003191** |
165
+ | Mean KLD | **0.014349** | **0.014023** | **0.014357** |
166
+ | Same top-p | **95.919%** | **95.516%** | **95.870%** |
167
+ | RMS Δp | **3.515%** | **3.679%** | **3.691%** |
 
 
 
 
 
168
 
169
+ IQ1_M is **8.633 MiB** smaller than IQ2_S and **36.523 MiB** smaller than Q3_K_M. IQ2_S retains the lowest Mean KLD of the three; IQ1_M has the smallest footprint, the lowest PPL ratio, and higher Same top-p than IQ2_S. The point-estimate differences remain small relative to the reported uncertainty scale.
170
 
171
  ## Frozen BF16-paired protocol
172
 
 
184
  | Batch / micro-batch | **512 / 512** |
185
  | Evaluator | `llama-perplexity` |
186
  | Reference logits | `--kl-divergence-base` |
187
+ | Quantized comparison | `--kl-divergence` |
188
  | BF16 logits reused across tiers | **Yes** |
189
  | llama.cpp revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` |
190
 
191
  ## Parent-model benchmark record
192
 
193
+ The scores below describe the **parent VeriLoop E2 release** and are not relabeled as quantization-specific reruns.
194
 
195
  | Benchmark | VeriLoop E2 parent score | Public evidence |
196
  |---|---:|---|
 
210
  # Overall sweet spot
211
  hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-Q6_K.gguf --local-dir .
212
 
213
+ # Minimum-footprint sweet spot
214
+ hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-IQ1_M.gguf --local-dir .
 
215
 
216
+ # Run
217
+ llama-server \
218
+ -m ./VeriLoop-E2-IQ1_M.gguf \
219
+ -ngl 99 \
220
+ -c 32768 \
221
+ --host 127.0.0.1 \
222
+ --port 8080
223
  ```
224
 
225
+ ## Memory guidance
226
 
227
+ | Variant | Main file | Positioning |
 
 
 
 
 
 
 
 
228
  |---|---:|---|
229
  | BF16 | **50.113 GiB** | Reference |
230
  | Q8_0 | **26.632 GiB** | High fidelity |
231
  | **Q6_K** | **20.566 GiB** | **Overall sweet spot** |
232
  | **Q5_K_M** | **18.965 GiB** | **Memory-quality sweet spot** |
233
+ | Q4_K_M | **18.301 GiB** | Balanced compact |
234
+ | Q3_K_M | **16.826 GiB** | Low-footprint alternative |
235
+ | IQ2_S | **16.799 GiB** | Lower-KLD low-footprint alternative |
236
+ | **IQ1_M** | **16.790 GiB** | **Minimum-footprint sweet spot** |
 
 
237
 
238
+ A **16.79 GiB model file does not imply full offload on a 16 GiB GPU**. Runtime memory also includes KV cache, compute buffers, allocator overhead, and optional speculative-decoding weights.
239
 
240
+ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128 41.999677 tok/s**), but there is no frozen paired BF16 throughput campaign. No universal speedup percentage is claimed for the quantization ladder.
 
 
 
 
 
 
 
 
 
 
 
241
 
242
  ## Measurement boundaries
243
 
244
+ - PPL, KLD, Same top-p, and token-probability statistics are **quantization-retention measurements**, not universal downstream capability-loss percentages.
245
+ - IQ1_M and IQ2_S are mixed-precision artifacts; the tier name does not equal the model-wide effective bit width.
 
 
 
 
246
  - Native 262K context comes from the parent configuration; practical context depends on runtime memory.
247
+ - MTP acceptance is prompt- and workload-dependent.
248
  - The model can still produce incorrect code, mathematics, scientific reasoning, or commands.
249
 
250
  ## Source model and evidence
 
254
  | Parent model | [VeriLoop E2](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2) |
255
  | GGUF repository | [VeriLoop E2 GGUF](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF) |
256
  | Quantization quality record | [Quantization Quality](./QUANTIZATION_QUALITY.md) |
257
+ | Artifact manifest | [Release Manifest](./RELEASE_MANIFEST.json) |
258
  | Technical report | [OpenReview](https://openreview.net/forum?id=P6FIQILHwX&noteId=P6FIQILHwX) |
259
  | Evaluation evidence | [VeriLoop E2 Evaluation Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence) |
260
  | Riemann ζ artifact | [Public artifact](https://github.com/brucewang123456789/GeniusTrail/tree/VeriLoop-E2/riemann-hypothesis) |
 
277
  }
278
  ```
279
 
280
+ For quantization-specific comparisons, identify the exact GGUF filename and corresponding manifest identity.
RELEASE_MANIFEST.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "schema": "veriloop.e2.gguf.release_manifest.v8_q1_local_identity_frozen",
3
  "repository": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF",
4
  "model": {
5
  "name": "VeriLoop E2",
@@ -16,23 +16,18 @@
16
  "quantization_methodology": "direct_from_BF16_no_low_bit_requantization",
17
  "quality_method": "paired_BF16_vs_quant_fixed_protocol",
18
  "parent_benchmark_rerun_per_quant": false,
19
- "performance_claim_boundary": "PPL_KLD_token_drift_are_quantization_fidelity_metrics_not_universal_task_score_loss",
20
- "immutable_identity_required_before_public_release": true,
21
- "lower_bit_alone_is_not_a_recommendation_criterion": true,
22
- "q3_filename_semantics": "custom_mixed_precision_release_tier_minimum_active_precision_Q3_K_not_uniform_3bit",
23
- "q2_filename_semantics": "custom_mixed_precision_release_tier_minimum_active_precision_IQ2_S_not_uniform_2bit",
24
- "q1_filename_semantics": "custom_mixed_precision_release_tier_minimum_active_precision_IQ1_M_not_uniform_1bit",
25
- "q1_prerelease_status": "QUANTITATIVE_PASS_LOCAL_IDENTITY_FROZEN_RUNTIME_MTP_REMOTE_PENDING",
26
- "sweet_spot_assignment_requires_full_release_validation": true
27
  },
28
- "sweet_spots": {
29
- "BF16": "reference_fidelity",
30
- "Q8_0": "high_fidelity_local",
31
- "Q6_K": "overall_quality_efficiency_sweet_spot",
32
- "Q5_K_M": "memory_quality_sweet_spot",
33
- "Q3_K_M": "adjacent_low_footprint_alternative",
34
- "IQ2_S": "current_released_low_footprint_sweet_spot",
35
- "IQ1_M": "ultra_low_footprint_frontier_candidate_pending_full_release_validation"
36
  },
37
  "quality_protocol": {
38
  "name": "paired_BF16_vs_quant",
@@ -48,78 +43,59 @@
48
  "micro_batch_size": 512,
49
  "evaluation_executable": "llama-perplexity",
50
  "reference_logit_method": "--kl-divergence-base",
51
- "candidate_comparison_method": "--kl-divergence",
52
  "corpus_sha256": "173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08",
53
  "bf16_logits_reused_across_quant_tiers": true,
54
  "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4"
55
  },
56
- "artifacts": {
57
  "VeriLoop-E2-BF16.gguf": {
58
  "tier": "BF16",
59
- "role": "canonical_reference_main",
60
  "bytes": 53808284000,
61
  "binary_gib": 50.11286959052086,
62
  "sha256": "11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d",
63
  "tensor_count": 851,
64
- "status": "VERIFIED"
65
  },
66
  "VeriLoop-E2-Q8_0.gguf": {
67
  "tier": "Q8_0",
68
- "role": "quantized_main",
69
  "bytes": 28595765600,
70
  "binary_gib": 26.631882041692734,
71
  "sha256": "6204a47274cfbc0c69c39877fb06615ce842bbab264eea77e2e0a5e3ae2fb8e8",
72
  "effective_bpw": 8.5,
73
- "status": "VERIFIED"
74
  },
75
  "VeriLoop-E2-Q6_K.gguf": {
76
  "tier": "Q6_K",
77
- "role": "quantized_main",
78
  "bytes": 22082532096,
79
  "binary_gib": 20.56596064567566,
80
  "sha256": "15d8f856471c4853f6bf0036b2a517426c6cb30a0313cef58a7f9577fd26fe9e",
81
  "effective_bpw": 6.57,
82
- "status": "VERIFIED"
83
  },
84
  "VeriLoop-E2-Q5_K_M.gguf": {
85
  "tier": "Q5_K_M",
86
- "role": "quantized_main",
87
  "bytes": 20363666176,
88
  "binary_gib": 18.965142011642456,
89
  "sha256": "f90ec14d7ec8f084a292413ae0e06483c87ab5e1422ddce51f2e3409d8853b61",
90
  "effective_bpw": 6.05,
91
- "status": "VERIFIED"
92
- },
93
- "VeriLoop-E2-Q4_K_M.gguf": {
94
- "tier": "Q4_K_M",
95
- "role": "quantized_main",
96
- "bytes": 19650634496,
97
- "binary_gib": 18.301079511642456,
98
- "sha256": null,
99
- "effective_bpw": 5.84,
100
- "status": "QUANTITATIVE_QC_PASS_RUNTIME_IDENTITY_PENDING"
101
  },
102
  "VeriLoop-E2-Q3_K_M.gguf": {
103
  "tier": "Q3_K_M",
104
- "role": "quantized_main",
105
  "bytes": 18066506496,
106
  "binary_gib": 16.825745344161987,
107
  "sha256": "c2e9539cfb85d87a99605b6914ed602c0fa750164c153ce5574710123aba06fe",
108
  "effective_bpw": 5.37,
109
- "status": "VERIFIED"
110
  },
111
  "VeriLoop-E2-IQ2_S.gguf": {
112
  "tier": "IQ2_S",
113
- "role": "quantized_main_minimum_footprint_sweet_spot",
114
- "quantization": "custom_mixed_IQ2_S_Q4_Q5_Q6",
115
- "source": "VeriLoop-E2-BF16.gguf",
116
- "direct_from_bf16": true,
117
- "low_bit_requantization": false,
118
  "bytes": 18037261056,
119
  "decimal_gb": 18.037261056,
120
  "binary_gib": 16.798508405685425,
121
  "sha256": "0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b",
122
- "sha256_state": "FROZEN",
123
  "tensor_count": 851,
124
  "tensor_type_counts": {
125
  "F32": 353,
@@ -128,33 +104,19 @@
128
  "Q5_K": 429,
129
  "Q6_K": 2
130
  },
131
- "effective_bpw": 5.36,
132
- "quantized_size_mib_reported_by_quantizer": 17191.19,
133
- "quantization_time_seconds": 168.08618,
134
- "importance_matrix_sha256": "72d7c3fb0ee839cb591deaa10201fb8e826f1f6b58a6a889f897b02f6381748b",
135
- "structure_gate": "PASS",
136
- "precision_policy_gate": "PASS",
137
- "quantitative_quality_gate": "PASS",
138
  "engineering_reserve_gate": "PASS",
139
  "runtime_gate": "PASS",
140
- "mtp_gate": "PASS",
141
- "remote_exact_size_gate": "PASS",
142
- "remote_full_sha256_gate": "PASS",
143
- "remote_readability_gate": "PASS",
144
- "status": "VERIFIED"
145
  },
146
  "VeriLoop-E2-IQ1_M.gguf": {
147
  "tier": "IQ1_M",
148
- "role": "ultra_low_footprint_frontier_candidate",
149
- "quantization": "custom_mixed_IQ1_M_IQ2_S_Q4_Q5_Q6",
150
- "source": "VeriLoop-E2-BF16.gguf",
151
- "direct_from_bf16": true,
152
- "low_bit_requantization": false,
153
  "bytes": 18028208896,
154
  "decimal_gb": 18.028208896,
155
  "binary_gib": 16.790077924728394,
156
  "sha256": "e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b",
157
- "sha256_state": "FROZEN_LOCAL",
158
  "tensor_count": 851,
159
  "tensor_type_counts": {
160
  "F32": 353,
@@ -164,23 +126,28 @@
164
  "Q5_K": 429,
165
  "Q6_K": 2
166
  },
167
- "effective_bpw": 5.36,
168
- "quantized_size_mib_reported_by_quantizer_v2": 17182.55,
169
- "structure_gate_v2": "PASS",
170
- "quantitative_quality_gate_v2": "PASS",
171
- "engineering_reserve_gate_v2": "PASS",
172
- "runtime_gate": "PENDING",
173
- "mtp_gate": "PENDING",
174
- "remote_release_gate": "PENDING",
175
- "status": "QUANTITATIVE_PASS_LOCAL_IDENTITY_FROZEN_RUNTIME_MTP_REMOTE_PENDING",
176
  "quantization_time_seconds": 167.77645,
177
  "end_to_end_quantization_seconds": 168.2609,
178
- "local_identity_gate": "PASS"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
179
  }
180
  },
181
  "quality_results": {
182
  "BF16": {
183
- "role": "reference",
184
  "mean_ppl": 4.840423,
185
  "ppl_uncertainty": 0.119931
186
  },
@@ -189,15 +156,12 @@
189
  "ppl_uncertainty": 0.120062,
190
  "ppl_ratio": 1.000643,
191
  "ppl_ratio_uncertainty": 0.000754,
192
- "relative_ppl_change_percent_observed": 0.0643,
193
  "mean_kld": 0.002176,
194
  "mean_kld_uncertainty": 0.000668,
195
  "same_top_p_percent": 98.815,
196
  "same_top_p_uncertainty_percent": 0.12,
197
  "rms_delta_p_percent": 1.251,
198
- "rms_delta_p_uncertainty_percent": 0.137,
199
  "log_ppl_correlation_percent": 99.95,
200
- "size_reduction_vs_bf16_percent": 46.856202290338786,
201
  "quantitative_quality_gate": "PASS"
202
  },
203
  "Q6_K": {
@@ -205,15 +169,12 @@
205
  "ppl_uncertainty": 0.119694,
206
  "ppl_ratio": 0.999605,
207
  "ppl_ratio_uncertainty": 0.001222,
208
- "relative_ppl_change_percent_observed": -0.0395,
209
  "mean_kld": 0.004409,
210
  "mean_kld_uncertainty": 0.000953,
211
  "same_top_p_percent": 98.204,
212
  "same_top_p_uncertainty_percent": 0.147,
213
  "rms_delta_p_percent": 1.929,
214
- "rms_delta_p_uncertainty_percent": 0.201,
215
  "log_ppl_correlation_percent": 99.88,
216
- "size_reduction_vs_bf16_percent": 58.960720442227824,
217
  "quantitative_quality_gate": "PASS"
218
  },
219
  "Q5_K_M": {
@@ -221,15 +182,12 @@
221
  "ppl_uncertainty": 0.12063,
222
  "ppl_ratio": 1.00445,
223
  "ppl_ratio_uncertainty": 0.001361,
224
- "relative_ppl_change_percent_observed": 0.445,
225
  "mean_kld": 0.006919,
226
  "mean_kld_uncertainty": 0.000945,
227
  "same_top_p_percent": 97.251,
228
  "same_top_p_uncertainty_percent": 0.181,
229
  "rms_delta_p_percent": 2.49,
230
- "rms_delta_p_uncertainty_percent": 0.211,
231
  "log_ppl_correlation_percent": 99.85,
232
- "size_reduction_vs_bf16_percent": 62.1551466387592,
233
  "quantitative_quality_gate": "PASS"
234
  },
235
  "Q4_K_M": {
@@ -237,15 +195,12 @@
237
  "ppl_uncertainty": 0.12072,
238
  "ppl_ratio": 1.004821,
239
  "ppl_ratio_uncertainty": 0.001722,
240
- "relative_ppl_change_percent_observed": 0.4821,
241
  "mean_kld": 0.0097,
242
  "mean_kld_uncertainty": 0.00111,
243
  "same_top_p_percent": 96.786,
244
  "same_top_p_uncertainty_percent": 0.195,
245
  "rms_delta_p_percent": 2.843,
246
- "rms_delta_p_uncertainty_percent": 0.155,
247
  "log_ppl_correlation_percent": 99.76,
248
- "size_reduction_vs_bf16_percent": 63.48028029290062,
249
  "quantitative_quality_gate": "PASS"
250
  },
251
  "Q3_K_M": {
@@ -253,15 +208,12 @@
253
  "ppl_uncertainty": 0.120606,
254
  "ppl_ratio": 1.00409,
255
  "ppl_ratio_uncertainty": 0.002181,
256
- "relative_ppl_change_percent_observed": 0.409,
257
  "mean_kld": 0.014349,
258
  "mean_kld_uncertainty": 0.001985,
259
  "same_top_p_percent": 95.919,
260
  "same_top_p_uncertainty_percent": 0.219,
261
  "rms_delta_p_percent": 3.515,
262
- "rms_delta_p_uncertainty_percent": 0.228,
263
  "log_ppl_correlation_percent": 99.62,
264
- "size_reduction_vs_bf16_percent": 66.4243028155293,
265
  "quantitative_quality_gate": "PASS"
266
  },
267
  "IQ2_S": {
@@ -269,39 +221,22 @@
269
  "ppl_uncertainty": 0.120482,
270
  "ppl_ratio": 1.003457,
271
  "ppl_ratio_uncertainty": 0.0021,
272
- "relative_ppl_change_percent_observed": 0.3457,
273
  "mean_kld": 0.014023,
274
  "mean_kld_uncertainty": 0.001408,
275
  "same_top_p_percent": 95.516,
276
  "same_top_p_uncertainty_percent": 0.229,
277
  "rms_delta_p_percent": 3.679,
278
- "rms_delta_p_uncertainty_percent": 0.231,
279
  "log_ppl_correlation_percent": 99.64,
280
- "size_reduction_vs_bf16_percent": 66.47865400056244,
281
  "quantitative_quality_gate": "PASS",
282
- "absolute_ppl_delta": 0.016736,
283
- "absolute_ppl_delta_uncertainty": 0.01016,
284
- "median_kld": 0.003619,
285
- "kld_p90": 0.022931,
286
- "kld_p95": 0.03865,
287
- "kld_p99": 0.129293,
288
- "kld_p999": 0.971304,
289
- "max_kld": 8.040164,
290
- "mean_delta_p_percent": -0.067,
291
- "mean_delta_p_uncertainty_percent": 0.041,
292
- "size_reduction_vs_q3_percent": 0.16187656427364672,
293
- "size_reduction_vs_q5_percent": 11.424294131976248,
294
  "engineering_reserve_gate": "PASS"
295
  },
296
  "IQ1_M": {
297
- "measurement_source": "V2_same_tensor_policy_as_V3",
298
  "mean_ppl": 4.85587,
299
  "ppl_uncertainty": 0.12048,
300
  "absolute_ppl_delta": 0.015446,
301
  "absolute_ppl_delta_uncertainty": 0.010522,
302
  "ppl_ratio": 1.003191,
303
  "ppl_ratio_uncertainty": 0.002175,
304
- "relative_ppl_change_percent_observed": 0.3191,
305
  "mean_kld": 0.014357,
306
  "mean_kld_uncertainty": 0.001317,
307
  "median_kld": 0.003778,
@@ -324,318 +259,47 @@
324
  "engineering_reserve_gate": "PASS"
325
  }
326
  },
327
- "release_gates": {
328
- "IQ2_S": {
329
- "ppl_ratio": {
330
- "hard_threshold_max": 1.015,
331
- "reserve_threshold_max": 1.0135,
332
- "measured": 1.003457,
333
- "hard_result": "PASS",
334
- "reserve_result": "PASS"
335
- },
336
- "mean_kld": {
337
- "hard_threshold_max": 0.015,
338
- "reserve_threshold_max": 0.0145,
339
- "measured": 0.014023,
340
- "hard_result": "PASS",
341
- "reserve_result": "PASS"
342
- },
343
- "same_top_p_percent": {
344
- "hard_threshold_min": 95.0,
345
- "reserve_threshold_min": 95.5,
346
- "measured": 95.516,
347
- "hard_result": "PASS",
348
- "reserve_result": "PASS"
349
- },
350
- "mixed_precision_policy": {
351
- "expected_counts": {
352
- "F32": 353,
353
- "IQ2_S": 3,
354
- "Q4_K": 64,
355
- "Q5_K": 429,
356
- "Q6_K": 2
357
- },
358
- "mismatch_count": 0,
359
- "result": "PASS"
360
- },
361
- "quantitative_quality_gate": "PASS",
362
- "engineering_reserve_gate": "PASS",
363
- "main_runtime_generation": "PASS",
364
- "mtp_runtime_load": "PASS",
365
- "mtp_runtime_generation": "PASS",
366
- "mtp_draft_model_load_evidence": "PASS",
367
- "mtp_draft_tokens_generated_gt_zero": "PASS",
368
- "mtp_draft_tokens_accepted_gt_zero": "PASS",
369
- "output_identity_audit": "DIVERGED_ADVISORY_NON_BLOCKING",
370
- "final_release_identity": "FROZEN",
371
- "remote_release_verification": "PASS",
372
- "overall_release_quality_gate": "PASS"
373
- },
374
- "IQ1_M": {
375
- "measurement_source": "V2_same_tensor_policy_as_V3",
376
- "ppl_ratio": {
377
- "hard_threshold_max": 1.015,
378
- "reserve_threshold_max": 1.0135,
379
- "measured": 1.003191,
380
- "hard_result": "PASS",
381
- "reserve_result": "PASS"
382
- },
383
- "mean_kld": {
384
- "hard_threshold_max": 0.015,
385
- "reserve_threshold_max": 0.0145,
386
- "measured": 0.014357,
387
- "hard_result": "PASS",
388
- "reserve_result": "PASS"
389
- },
390
- "same_top_p_percent": {
391
- "hard_threshold_min": 95.0,
392
- "reserve_threshold_min": 95.5,
393
- "measured": 95.87,
394
- "hard_result": "PASS",
395
- "reserve_result": "PASS"
396
- },
397
- "mixed_precision_policy": {
398
- "expected_counts": {
399
- "F32": 353,
400
- "IQ1_M": 1,
401
- "IQ2_S": 2,
402
- "Q4_K": 64,
403
- "Q5_K": 429,
404
- "Q6_K": 2
405
- },
406
- "mismatch_count": 0,
407
- "result": "PASS"
408
- },
409
- "quantitative_quality_gate": "PASS",
410
- "engineering_reserve_gate": "PASS",
411
- "obsolete_v2_point_estimate_dominance_gate": "REMOVED_IN_V3",
412
- "stock_runtime_gate": "PENDING",
413
- "mtp_gate": "PENDING",
414
- "remote_release_gate": "PENDING",
415
- "overall_release_quality_gate": "PENDING_RUNTIME_MTP_REMOTE",
416
- "final_local_identity_gate": "PASS"
417
- }
418
- },
419
  "runtime_validation": {
420
  "IQ2_S": {
421
  "stock_llama_cpp_only": true,
422
- "veriloop_harness_used": false,
423
  "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4",
424
- "main_model_load": "PASS",
425
- "main_model_generation": "PASS",
426
- "mtp_interface": "PASS",
427
- "mtp_load": "PASS",
428
- "mtp_generation": "PASS",
429
- "mtp_draft_model_load_evidence": "PASS",
430
  "draft_tokens_generated": 331,
431
  "draft_tokens_accepted": 116,
432
- "draft_acceptance_rate": 0.35045,
433
- "mean_draft_length": 3.76,
434
- "output_identity_audit": "DIVERGED",
435
- "output_identity_is_release_blocker": false,
436
- "mtp_bytes": 2008056096,
437
- "mtp_sha256": "0e7f2dfe254f3a5d195d105832411acc1007fd7bb9ac980fe9aae4a1f290d671",
438
- "runtime_mtp_gate": "PASS"
439
  },
440
  "IQ1_M": {
441
- "state": "PENDING_V3",
442
  "stock_llama_cpp_only": true,
443
- "required": [
444
- "main_model_load",
445
- "main_model_generation",
446
- "mtp_load",
447
- "mtp_generation",
448
- "mtp_draft_engagement"
449
- ],
450
- "claim": "NO_RUNTIME_OR_MTP_PASS_CLAIM_BEFORE_V3_EVIDENCE"
451
- }
452
- },
453
- "remote_repository_verification": {
454
- "IQ2_S": {
455
- "state": "PASS",
456
- "remote_path": "VeriLoop-E2-IQ2_S.gguf",
457
- "main_remote_bytes": 18037261056,
458
- "main_remote_lfs_sha256": "0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b",
459
- "remote_exact_size_gate": "PASS",
460
- "remote_full_sha256_gate": "PASS",
461
- "remote_readability_gate": "PASS",
462
- "remote_head_identity_gate": "PASS",
463
- "remote_tail_identity_gate": "PASS",
464
- "remote_release_gate": "PASS"
465
- },
466
- "IQ1_M": {
467
- "state": "PENDING_V3",
468
- "remote_path": "VeriLoop-E2-IQ1_M.gguf",
469
- "required": [
470
- "remote_exact_size",
471
- "remote_full_sha256",
472
- "remote_readability",
473
- "remote_head_identity",
474
- "remote_tail_identity"
475
- ]
476
- }
477
- },
478
- "throughput": {
479
- "IQ2_S": {
480
- "paired_speed_protocol_run": false,
481
- "speedup_claimed": false,
482
- "mtp_draft_tokens_generated": 331,
483
- "mtp_draft_tokens_accepted": 116,
484
- "mtp_draft_acceptance_rate": 0.35045,
485
- "note": "Draft acceptance is reported descriptively; no universal speedup is claimed without a frozen paired throughput protocol."
486
- },
487
- "IQ1_M": {
488
- "paired_speed_protocol_run": false,
489
- "speedup_claimed": false,
490
- "mtp_acceptance_measured": false,
491
- "note": "No IQ1_M throughput or MTP acceptance claim before final V3 runtime validation."
492
  }
493
  },
494
- "release_matrix": {
495
- "BF16": {
496
- "state": "released",
497
- "role": "canonical_reference"
498
- },
499
- "Q8_0": {
500
- "state": "released",
501
- "role": "high_fidelity_local"
502
- },
503
- "Q6_K": {
504
- "state": "released",
505
- "role": "overall_quality_efficiency_sweet_spot",
506
- "quality_record": "PASS"
507
- },
508
- "Q5_K_M": {
509
- "state": "released",
510
- "role": "memory_quality_sweet_spot",
511
- "quality_record": "PASS"
512
- },
513
- "Q4_K_M": {
514
- "state": "release_candidate",
515
- "role": "intermediate_mixed_precision_candidate",
516
- "quality_record": "PASS",
517
- "runtime_record": "PENDING"
518
- },
519
- "Q3_K_M": {
520
- "state": "verified_release",
521
- "role": "adjacent_low_footprint_alternative",
522
- "quality_record": "PASS",
523
- "runtime_record": "PASS",
524
- "remote_release": "PASS"
525
- },
526
- "IQ2_S": {
527
- "state": "verified_release",
528
- "role": "minimum_footprint_low_footprint_sweet_spot",
529
- "quality_record": "PASS",
530
- "engineering_reserve_record": "PASS",
531
- "runtime_record": "PASS",
532
- "mtp_record": "PASS",
533
- "release_identity": "FROZEN",
534
- "remote_release": "PASS",
535
- "overall_release": "PASS"
536
- },
537
  "IQ1_M": {
538
- "state": "prerelease_quantitative_pass_local_identity_frozen",
539
- "role": "ultra_low_footprint_frontier_candidate",
540
- "quality_record": "PASS_V2_SAME_POLICY",
541
- "engineering_reserve_record": "PASS_V2_SAME_POLICY",
542
- "runtime_record": "PENDING",
543
- "mtp_record": "PENDING",
544
- "release_identity": "FROZEN_LOCAL",
545
- "remote_release": "PENDING",
546
- "overall_release": "PENDING_RUNTIME_MTP_REMOTE"
547
- }
548
- },
549
- "current_release_state": {
550
- "BF16": "RELEASED_REFERENCE",
551
- "Q8_0": "RELEASED_VERIFIED",
552
- "Q6_K": "RELEASED_VERIFIED_OVERALL_SWEET_SPOT",
553
- "Q5_K_M": "RELEASED_VERIFIED_MEMORY_QUALITY_SWEET_SPOT",
554
- "Q4_K_M": "QUANTITATIVE_QC_PASS_RUNTIME_AND_SHA_PENDING",
555
- "Q3_K_M": "VERIFIED_RELEASE_LOW_FOOTPRINT_ALTERNATIVE",
556
- "IQ2_S": "VERIFIED_RELEASE_CURRENT_LOW_FOOTPRINT_SWEET_SPOT",
557
- "IQ1_M": "PRERELEASE_QUANTITATIVE_PASS_LOCAL_IDENTITY_FROZEN_RUNTIME_MTP_REMOTE_PENDING"
558
- },
559
- "deployment_recommendation": {
560
- "overall_sweet_spot": "Q6_K",
561
- "memory_quality_sweet_spot": "Q5_K_M",
562
- "current_released_low_footprint_sweet_spot": "IQ2_S",
563
- "ultra_low_footprint_frontier_candidate": "IQ1_M",
564
- "q1_post_validation_promotion_rule": "If the identical V3 policy reproduces the V2 hard/reserve fidelity record and passes immutable identity, stock runtime, MTP, and remote verification, promote IQ1_M to minimum-footprint sweet spot; retain IQ2_S as the lower-KLD low-footprint alternative.",
565
- "basis": "IQ1_M is 16.790078 GiB versus IQ2_S at 16.798508 GiB. IQ1_M has a lower PPL ratio (1.003191 vs 1.003457) and higher Same top-p (95.870% vs 95.516%), while IQ2_S has lower Mean KLD (0.014023 vs 0.014357) and slightly lower RMS delta-p. The point-estimate differences are small relative to reported uncertainty, so the two are treated as neighboring Pareto points rather than categorical quality winners. IQ2_S retains the current released low-footprint sweet-spot designation because its immutable identity, stock runtime, MTP, and remote release verification are already closed."
566
- },
567
- "q2_development": {
568
- "v1_rejected": {
569
- "target": "four_low_sensitivity_ffn_up_tensors_to_IQ2_S",
570
- "bytes": 18049098496,
571
- "binary_gib": 16.809533,
572
- "ppl_ratio": 1.008995,
573
- "mean_kld": 0.014951,
574
- "same_top_p_percent": 95.809,
575
- "hard_gate": "PASS",
576
- "engineering_reserve_gate": "FAIL",
577
- "failure_axis": "MEAN_KLD",
578
- "large_candidate_retention": "DELETE_FAILED_ARTIFACT_PRESERVE_REPORTS"
579
- },
580
- "v2_final": {
581
- "target": "three_q3_spine_ffn_down_tensors_to_IQ2_S",
582
- "exact_tensors": [
583
- "blk.0.ffn_down.weight",
584
- "blk.1.ffn_down.weight",
585
- "blk.3.ffn_down.weight"
586
- ],
587
- "bytes": 18037261056,
588
- "binary_gib": 16.798508405685425,
589
- "ppl_ratio": 1.003457,
590
- "mean_kld": 0.014023,
591
- "same_top_p_percent": 95.516,
592
- "hard_gate": "PASS",
593
- "engineering_reserve_gate": "PASS"
594
- }
595
- },
596
- "q1_development": {
597
- "v1_rejected": {
598
- "policy": "blk.1_and_blk.0_ffn_down_to_IQ1_M_blk.3_restored_Q3_K",
599
- "bytes": 18028905216,
600
- "binary_gib": 16.790726,
601
- "ppl_ratio": 1.006006,
602
- "mean_kld": 0.016349,
603
- "same_top_p_percent": 95.736,
604
- "hard_quality_gate": "FAIL",
605
- "engineering_reserve_gate": "FAIL",
606
- "candidate_file_deleted": true
607
- },
608
- "v2_measured": {
609
- "policy": "single_factor_blk.1_ffn_down_IQ2_S_to_IQ1_M_everything_else_identical_to_Q2",
610
- "bytes": 18028208896,
611
- "binary_gib": 16.790077924728394,
612
- "ppl_ratio": 1.003191,
613
- "mean_kld": 0.014357,
614
- "same_top_p_percent": 95.87,
615
- "hard_quality_gate": "PASS",
616
- "engineering_reserve_gate": "PASS",
617
- "obsolete_point_estimate_dominance_gate": "FAIL_MEAN_KLD_VS_Q2",
618
- "candidate_file_deleted": true,
619
- "interpretation": "Quantitative release quality passed; file deletion was caused only by an experimental cross-tier dominance rule removed in V3."
620
- },
621
- "accepted_rebuild": {
622
- "quantization_policy_change_vs_v2": false,
623
- "hard_gate_change_vs_v2": false,
624
- "reserve_gate_change_vs_v2": false,
625
- "bytes": 18028208896,
626
- "binary_gib": 16.790078,
627
- "ppl_ratio": 1.003191,
628
- "mean_kld": 0.014357,
629
- "same_top_p_percent": 95.87,
630
  "hard_quality_gate": "PASS",
631
  "engineering_reserve_gate": "PASS",
632
- "sweetspot_promotion_gate": "PASS",
633
- "candidate_acceptance_gate": "PASS",
634
- "sha256": "e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b",
635
- "local_identity": "FROZEN",
636
- "stock_runtime": "PENDING",
637
- "mtp": "PENDING",
638
- "remote_release": "PENDING"
639
  }
 
 
 
 
640
  }
641
  }
 
1
  {
2
+ "schema": "veriloop.e2.gguf.release_manifest.v9",
3
  "repository": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF",
4
  "model": {
5
  "name": "VeriLoop E2",
 
16
  "quantization_methodology": "direct_from_BF16_no_low_bit_requantization",
17
  "quality_method": "paired_BF16_vs_quant_fixed_protocol",
18
  "parent_benchmark_rerun_per_quant": false,
19
+ "performance_claim_boundary": "PPL_KLD_and_token_drift_are_quantization_fidelity_metrics_not_universal_task_score_loss",
20
+ "q3_filename_semantics": "mixed_precision_tier_minimum_active_precision_Q3_K",
21
+ "q2_filename_semantics": "mixed_precision_tier_minimum_active_precision_IQ2_S",
22
+ "q1_filename_semantics": "mixed_precision_tier_minimum_active_precision_IQ1_M",
23
+ "internal_release_readiness": "PASS"
 
 
 
24
  },
25
+ "recommended_variants": {
26
+ "overall_sweet_spot": "Q6_K",
27
+ "memory_quality_sweet_spot": "Q5_K_M",
28
+ "minimum_footprint_sweet_spot": "IQ1_M",
29
+ "lower_kld_low_footprint_alternative": "IQ2_S",
30
+ "adjacent_low_footprint_alternative": "Q3_K_M"
 
 
31
  },
32
  "quality_protocol": {
33
  "name": "paired_BF16_vs_quant",
 
43
  "micro_batch_size": 512,
44
  "evaluation_executable": "llama-perplexity",
45
  "reference_logit_method": "--kl-divergence-base",
46
+ "quantized_comparison_method": "--kl-divergence",
47
  "corpus_sha256": "173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08",
48
  "bf16_logits_reused_across_quant_tiers": true,
49
  "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4"
50
  },
51
+ "frozen_artifacts": {
52
  "VeriLoop-E2-BF16.gguf": {
53
  "tier": "BF16",
 
54
  "bytes": 53808284000,
55
  "binary_gib": 50.11286959052086,
56
  "sha256": "11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d",
57
  "tensor_count": 851,
58
+ "identity_gate": "PASS"
59
  },
60
  "VeriLoop-E2-Q8_0.gguf": {
61
  "tier": "Q8_0",
 
62
  "bytes": 28595765600,
63
  "binary_gib": 26.631882041692734,
64
  "sha256": "6204a47274cfbc0c69c39877fb06615ce842bbab264eea77e2e0a5e3ae2fb8e8",
65
  "effective_bpw": 8.5,
66
+ "identity_gate": "PASS"
67
  },
68
  "VeriLoop-E2-Q6_K.gguf": {
69
  "tier": "Q6_K",
 
70
  "bytes": 22082532096,
71
  "binary_gib": 20.56596064567566,
72
  "sha256": "15d8f856471c4853f6bf0036b2a517426c6cb30a0313cef58a7f9577fd26fe9e",
73
  "effective_bpw": 6.57,
74
+ "identity_gate": "PASS"
75
  },
76
  "VeriLoop-E2-Q5_K_M.gguf": {
77
  "tier": "Q5_K_M",
 
78
  "bytes": 20363666176,
79
  "binary_gib": 18.965142011642456,
80
  "sha256": "f90ec14d7ec8f084a292413ae0e06483c87ab5e1422ddce51f2e3409d8853b61",
81
  "effective_bpw": 6.05,
82
+ "identity_gate": "PASS"
 
 
 
 
 
 
 
 
 
83
  },
84
  "VeriLoop-E2-Q3_K_M.gguf": {
85
  "tier": "Q3_K_M",
 
86
  "bytes": 18066506496,
87
  "binary_gib": 16.825745344161987,
88
  "sha256": "c2e9539cfb85d87a99605b6914ed602c0fa750164c153ce5574710123aba06fe",
89
  "effective_bpw": 5.37,
90
+ "identity_gate": "PASS"
91
  },
92
  "VeriLoop-E2-IQ2_S.gguf": {
93
  "tier": "IQ2_S",
 
 
 
 
 
94
  "bytes": 18037261056,
95
  "decimal_gb": 18.037261056,
96
  "binary_gib": 16.798508405685425,
97
  "sha256": "0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b",
98
+ "effective_bpw": 5.36,
99
  "tensor_count": 851,
100
  "tensor_type_counts": {
101
  "F32": 353,
 
104
  "Q5_K": 429,
105
  "Q6_K": 2
106
  },
107
+ "identity_gate": "PASS",
108
+ "quality_gate": "PASS",
 
 
 
 
 
109
  "engineering_reserve_gate": "PASS",
110
  "runtime_gate": "PASS",
111
+ "mtp_gate": "PASS"
 
 
 
 
112
  },
113
  "VeriLoop-E2-IQ1_M.gguf": {
114
  "tier": "IQ1_M",
 
 
 
 
 
115
  "bytes": 18028208896,
116
  "decimal_gb": 18.028208896,
117
  "binary_gib": 16.790077924728394,
118
  "sha256": "e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b",
119
+ "effective_bpw": 5.36,
120
  "tensor_count": 851,
121
  "tensor_type_counts": {
122
  "F32": 353,
 
126
  "Q5_K": 429,
127
  "Q6_K": 2
128
  },
129
+ "quantized_size_mib_reported_by_quantizer": 17182.55,
 
 
 
 
 
 
 
 
130
  "quantization_time_seconds": 167.77645,
131
  "end_to_end_quantization_seconds": 168.2609,
132
+ "identity_gate": "PASS",
133
+ "structure_gate": "PASS",
134
+ "quality_gate": "PASS",
135
+ "engineering_reserve_gate": "PASS",
136
+ "runtime_gate": "PASS",
137
+ "mtp_gate": "PASS",
138
+ "release_readiness_gate": "PASS"
139
+ }
140
+ },
141
+ "measured_tiers": {
142
+ "Q4_K_M": {
143
+ "bytes": 19650634496,
144
+ "binary_gib": 18.301079511642456,
145
+ "effective_bpw": 5.84,
146
+ "quantitative_quality_gate": "PASS"
147
  }
148
  },
149
  "quality_results": {
150
  "BF16": {
 
151
  "mean_ppl": 4.840423,
152
  "ppl_uncertainty": 0.119931
153
  },
 
156
  "ppl_uncertainty": 0.120062,
157
  "ppl_ratio": 1.000643,
158
  "ppl_ratio_uncertainty": 0.000754,
 
159
  "mean_kld": 0.002176,
160
  "mean_kld_uncertainty": 0.000668,
161
  "same_top_p_percent": 98.815,
162
  "same_top_p_uncertainty_percent": 0.12,
163
  "rms_delta_p_percent": 1.251,
 
164
  "log_ppl_correlation_percent": 99.95,
 
165
  "quantitative_quality_gate": "PASS"
166
  },
167
  "Q6_K": {
 
169
  "ppl_uncertainty": 0.119694,
170
  "ppl_ratio": 0.999605,
171
  "ppl_ratio_uncertainty": 0.001222,
 
172
  "mean_kld": 0.004409,
173
  "mean_kld_uncertainty": 0.000953,
174
  "same_top_p_percent": 98.204,
175
  "same_top_p_uncertainty_percent": 0.147,
176
  "rms_delta_p_percent": 1.929,
 
177
  "log_ppl_correlation_percent": 99.88,
 
178
  "quantitative_quality_gate": "PASS"
179
  },
180
  "Q5_K_M": {
 
182
  "ppl_uncertainty": 0.12063,
183
  "ppl_ratio": 1.00445,
184
  "ppl_ratio_uncertainty": 0.001361,
 
185
  "mean_kld": 0.006919,
186
  "mean_kld_uncertainty": 0.000945,
187
  "same_top_p_percent": 97.251,
188
  "same_top_p_uncertainty_percent": 0.181,
189
  "rms_delta_p_percent": 2.49,
 
190
  "log_ppl_correlation_percent": 99.85,
 
191
  "quantitative_quality_gate": "PASS"
192
  },
193
  "Q4_K_M": {
 
195
  "ppl_uncertainty": 0.12072,
196
  "ppl_ratio": 1.004821,
197
  "ppl_ratio_uncertainty": 0.001722,
 
198
  "mean_kld": 0.0097,
199
  "mean_kld_uncertainty": 0.00111,
200
  "same_top_p_percent": 96.786,
201
  "same_top_p_uncertainty_percent": 0.195,
202
  "rms_delta_p_percent": 2.843,
 
203
  "log_ppl_correlation_percent": 99.76,
 
204
  "quantitative_quality_gate": "PASS"
205
  },
206
  "Q3_K_M": {
 
208
  "ppl_uncertainty": 0.120606,
209
  "ppl_ratio": 1.00409,
210
  "ppl_ratio_uncertainty": 0.002181,
 
211
  "mean_kld": 0.014349,
212
  "mean_kld_uncertainty": 0.001985,
213
  "same_top_p_percent": 95.919,
214
  "same_top_p_uncertainty_percent": 0.219,
215
  "rms_delta_p_percent": 3.515,
 
216
  "log_ppl_correlation_percent": 99.62,
 
217
  "quantitative_quality_gate": "PASS"
218
  },
219
  "IQ2_S": {
 
221
  "ppl_uncertainty": 0.120482,
222
  "ppl_ratio": 1.003457,
223
  "ppl_ratio_uncertainty": 0.0021,
 
224
  "mean_kld": 0.014023,
225
  "mean_kld_uncertainty": 0.001408,
226
  "same_top_p_percent": 95.516,
227
  "same_top_p_uncertainty_percent": 0.229,
228
  "rms_delta_p_percent": 3.679,
 
229
  "log_ppl_correlation_percent": 99.64,
 
230
  "quantitative_quality_gate": "PASS",
 
 
 
 
 
 
 
 
 
 
 
 
231
  "engineering_reserve_gate": "PASS"
232
  },
233
  "IQ1_M": {
 
234
  "mean_ppl": 4.85587,
235
  "ppl_uncertainty": 0.12048,
236
  "absolute_ppl_delta": 0.015446,
237
  "absolute_ppl_delta_uncertainty": 0.010522,
238
  "ppl_ratio": 1.003191,
239
  "ppl_ratio_uncertainty": 0.002175,
 
240
  "mean_kld": 0.014357,
241
  "mean_kld_uncertainty": 0.001317,
242
  "median_kld": 0.003778,
 
259
  "engineering_reserve_gate": "PASS"
260
  }
261
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
262
  "runtime_validation": {
263
  "IQ2_S": {
264
  "stock_llama_cpp_only": true,
 
265
  "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4",
266
+ "main_runtime": "PASS",
267
+ "mtp_runtime": "PASS",
268
+ "mtp_engagement": "PASS",
 
 
 
269
  "draft_tokens_generated": 331,
270
  "draft_tokens_accepted": 116,
271
+ "draft_acceptance_rate": 0.35045
 
 
 
 
 
 
272
  },
273
  "IQ1_M": {
 
274
  "stock_llama_cpp_only": true,
275
+ "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4",
276
+ "main_http_status": 200,
277
+ "main_runtime": "PASS",
278
+ "mtp_http_status": 200,
279
+ "mtp_runtime": "PASS",
280
+ "mtp_engagement": "PASS",
281
+ "draft_tokens_generated": 104,
282
+ "draft_tokens_accepted": 76,
283
+ "draft_acceptance_rate": 0.73076923,
284
+ "acceptance_rate_consistency_gate": "PASS",
285
+ "output_identity_audit": "IDENTICAL",
286
+ "runtime_mtp_gate": "PASS"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
287
  }
288
  },
289
+ "release_validation": {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
290
  "IQ1_M": {
291
+ "identity_gate": "PASS",
292
+ "structure_gate": "PASS",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
293
  "hard_quality_gate": "PASS",
294
  "engineering_reserve_gate": "PASS",
295
+ "stock_runtime_gate": "PASS",
296
+ "mtp_runtime_gate": "PASS",
297
+ "mtp_engagement_gate": "PASS",
298
+ "release_readiness_gate": "PASS"
 
 
 
299
  }
300
+ },
301
+ "throughput_claims": {
302
+ "paired_bf16_speedup_claimed": false,
303
+ "note": "MTP acceptance is prompt-dependent; no universal throughput speedup is claimed without a frozen paired throughput protocol."
304
  }
305
  }