ConorWang commited on
Commit
4763719
·
verified ·
1 Parent(s): 150097e

Upload README.md

Browse files
Files changed (1) hide show
  1. README.md +10 -10
README.md CHANGED
@@ -70,7 +70,7 @@ tags:
70
  | **Default** | **`VeriLoop-E2-Q6_K.gguf`** | **20.566 GiB** | BF16-paired PPL parity within uncertainty; KLD **0.004409**; Same top-p **98.204%** | Strongest overall quality / footprint balance |
71
  | **Tighter memory** | **`VeriLoop-E2-Q5_K_M.gguf`** | **18.965 GiB** | PPL **+0.4450%**; KLD **0.006919**; Same top-p **97.251%** | Smaller deployment with a wider fidelity margin |
72
  | **Released low-footprint sweet spot** | **`VeriLoop-E2-IQ2_S.gguf`** | **16.799 GiB** | PPL **+0.3457%**; KLD **0.014023**; Same top-p **95.516%** | Smallest fully release-verified low-footprint tier |
73
- | **Ultra-low-footprint frontier candidate** | **`VeriLoop-E2-IQ1_M.gguf`** | **16.790 GiB** | PPL **+0.3191%**; KLD **0.014357**; Same top-p **95.870%** | Smallest measured candidate; local immutable SHA256 frozen; stock runtime/MTP/remote verification pending |
74
  | **Low-footprint alternative** | `VeriLoop-E2-Q3_K_M.gguf` | **16.826 GiB** | PPL **+0.4090%**; KLD **0.014349**; Same top-p **95.919%** | Adjacent fully validated option |
75
  | **High fidelity** | `VeriLoop-E2-Q8_0.gguf` | **26.632 GiB** | PPL **+0.0643%**; KLD **0.002176**; Same top-p **98.815%** | Distributional fidelity matters more than footprint |
76
  | **Reference** | `VeriLoop-E2-BF16.gguf` | **50.113 GiB** | Canonical reference | Reproducibility, attribution, quantization research |
@@ -83,9 +83,9 @@ tags:
83
 
84
  **IQ2_S remains the current released low-footprint sweet spot.** It is fully verified through structure, BF16-paired fidelity, engineering reserve, stock llama.cpp runtime, MTP engagement, immutable SHA256, and remote readability.
85
 
86
- **IQ1_M is the current ultra-low-footprint frontier candidate, not yet a fully released replacement for IQ2_S.** Its measured V2 artifact is **16.790078 GiB**, **66.4955% smaller than BF16**, **8.633 MiB / 0.0502% smaller than IQ2_S**, and **36.523 MiB / 0.2120% smaller than Q3_K_M**. Against IQ2_S, IQ1_M has a slightly better PPL ratio (**1.003191 vs 1.003457**) and higher Same top-p (**95.870% vs 95.516%**), but a slightly higher Mean KLD (**0.014357 vs 0.014023**). The uncertainty bands substantially overlap, so this release does **not** claim that either low-footprint tier is universally superior in fidelity.
87
 
88
- **Promotion rule.** If the identical V3 rebuild reproduces these measurements and passes stock runtime, MTP, immutable identity, and remote verification, IQ1_M should be promoted to the **minimum-footprint sweet spot**. IQ2_S would remain the **released low-footprint balance point / lower-KLD alternative**. Until those final gates close, IQ2_S keeps the released sweet-spot designation.
89
 
90
  > **Naming note:** `VeriLoop-E2-IQ1_M.gguf` is not a uniform 1-bit model. Its measured policy is **353 F32 + 1 IQ1_M + 2 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K = 851 tensors**, with **5.36 effective BPW**. `VeriLoop-E2-IQ2_S.gguf` is likewise a mixed-precision release rather than a uniform 2-bit model.
91
 
@@ -97,10 +97,10 @@ tags:
97
  | **Q8_0** | High-fidelity local | **28.596 GB / 26.632 GiB** | **46.86%** | **8.50 BPW** | Verified |
98
  | **Q6_K** | **Overall sweet spot** | **22.083 GB / 20.566 GiB** | **58.96%** | **6.57 BPW** | Verified / Release Quality PASS |
99
  | **Q5_K_M** | **Memory-quality sweet spot** | **20.364 GB / 18.965 GiB** | **62.16%** | **6.05 BPW** | Verified / Release Quality PASS |
100
- | **Q4_K_M** | Intermediate mixed-precision candidate | **19.651 GB / 18.301 GiB** | **63.48%** | **5.84 BPW** | Quantitative QC PASS; own runtime/identity pending |
101
  | **Q3_K_M** | Adjacent low-footprint alternative | **18.067 GB / 16.826 GiB** | **66.42%** | **5.37 BPW** | Full release PASS |
102
  | **IQ2_S** | **Released low-footprint sweet spot** | **18.037 GB / 16.799 GiB** | **66.48%** | **5.36 BPW** | **Full release PASS** |
103
- | **IQ1_M** | **Ultra-low-footprint frontier candidate** | **18.028 GB / 16.790 GiB** | **66.50%** | **5.36 BPW** | **Structure + hard/reserve fidelity PASS; local identity frozen; runtime/MTP/remote pending** |
104
 
105
  ## Quantization-retention benchmark
106
 
@@ -290,7 +290,7 @@ hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-Q6_K.gguf --loc
290
  hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-IQ2_S.gguf --local-dir .
291
  ```
292
 
293
- > **IQ1_M is not listed as a download command in this pre-release draft.** Its quantitative evidence is PASS and its final local SHA256 is frozen as `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`, but stock runtime, MTP, upload, and remote identity verification must still be completed before public-release promotion.
294
 
295
  Run IQ2_S:
296
 
@@ -317,11 +317,11 @@ llama-server -m ./VeriLoop-E2-IQ2_S.gguf --model-draft ./mtp-VeriLoop-E2-Q5_
317
  | Q4_K_M | **18.301 GiB** | Intermediate quantitative candidate |
318
  | Q3_K_M | **16.826 GiB** | Adjacent low-footprint release |
319
  | **IQ2_S** | **16.799 GiB** | **Released low-footprint sweet spot** |
320
- | **IQ1_M** | **16.790 GiB** | **Ultra-low-footprint frontier candidate; local identity frozen; runtime/MTP/remote gates pending** |
321
 
322
  A **16.79 GiB model file does not imply full offload on a 16 GiB GPU**. Runtime memory also includes KV cache, compute buffers, allocator overhead, and optional MTP weights.
323
 
324
- A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128 41.999677 tok/s**), but there is no frozen paired BF16 throughput campaign. No universal speedup percentage is claimed for Q6/Q5/Q4/Q3/IQ2_S/IQ1_M. IQ1_M throughput and MTP acceptance remain **pending** until its final stock-runtime validation.
325
 
326
  ## Release positioning
327
 
@@ -334,7 +334,7 @@ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128
334
  | Q4_K_M | Intermediate mixed-precision candidate |
335
  | Q3_K_M | Adjacent low-footprint release |
336
  | **IQ2_S** | **Current released low-footprint sweet spot / lower-KLD frontier point** |
337
- | **IQ1_M** | **Ultra-low-footprint frontier candidate; promote to minimum-footprint sweet spot only after V3 runtime/MTP/remote PASS** |
338
 
339
  ## Measurement boundaries
340
 
@@ -342,7 +342,7 @@ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128
342
  - IQ1_M's **+0.3191% PPL** does not mean a 0.3191% drop on SWE-bench, Terminal-Bench, AIME, GPQA, or another downstream benchmark.
343
  - IQ2_S's **+0.3457% PPL** likewise does not imply a 0.3457% downstream-score loss.
344
  - IQ1_M and IQ2_S are both mixed-precision artifacts; the nominal tier name is not the model-wide effective bit width.
345
- - IQ1_M now has a frozen local SHA256 (`e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`) and quantitative PASS evidence, but **stock-runtime/MTP PASS and remote release identity are still pending**.
346
  - The IQ1_M vs IQ2_S point-estimate differences are small relative to the reported uncertainty; this repository does not claim categorical fidelity superiority for either frontier point.
347
  - Native 262K context comes from the parent configuration; practical context depends on runtime memory.
348
  - MTP acceptance is prompt/workload-dependent.
 
70
  | **Default** | **`VeriLoop-E2-Q6_K.gguf`** | **20.566 GiB** | BF16-paired PPL parity within uncertainty; KLD **0.004409**; Same top-p **98.204%** | Strongest overall quality / footprint balance |
71
  | **Tighter memory** | **`VeriLoop-E2-Q5_K_M.gguf`** | **18.965 GiB** | PPL **+0.4450%**; KLD **0.006919**; Same top-p **97.251%** | Smaller deployment with a wider fidelity margin |
72
  | **Released low-footprint sweet spot** | **`VeriLoop-E2-IQ2_S.gguf`** | **16.799 GiB** | PPL **+0.3457%**; KLD **0.014023**; Same top-p **95.516%** | Smallest fully release-verified low-footprint tier |
73
+ | **Ultra-low-footprint frontier candidate** | **`VeriLoop-E2-IQ1_M.gguf`** | **16.790 GiB** | PPL **+0.3191%**; KLD **0.014357**; Same top-p **95.870%** | Smallest measured candidate; quantitative fidelity and local identity verified |
74
  | **Low-footprint alternative** | `VeriLoop-E2-Q3_K_M.gguf` | **16.826 GiB** | PPL **+0.4090%**; KLD **0.014349**; Same top-p **95.919%** | Adjacent fully validated option |
75
  | **High fidelity** | `VeriLoop-E2-Q8_0.gguf` | **26.632 GiB** | PPL **+0.0643%**; KLD **0.002176**; Same top-p **98.815%** | Distributional fidelity matters more than footprint |
76
  | **Reference** | `VeriLoop-E2-BF16.gguf` | **50.113 GiB** | Canonical reference | Reproducibility, attribution, quantization research |
 
83
 
84
  **IQ2_S remains the current released low-footprint sweet spot.** It is fully verified through structure, BF16-paired fidelity, engineering reserve, stock llama.cpp runtime, MTP engagement, immutable SHA256, and remote readability.
85
 
86
+ **IQ1_M is the current ultra-low-footprint frontier candidate.** Its measured V2 artifact is **16.790078 GiB**, **66.4955% smaller than BF16**, **8.633 MiB / 0.0502% smaller than IQ2_S**, and **36.523 MiB / 0.2120% smaller than Q3_K_M**. Against IQ2_S, IQ1_M has a slightly better PPL ratio (**1.003191 vs 1.003457**) and higher Same top-p (**95.870% vs 95.516%**), but a slightly higher Mean KLD (**0.014357 vs 0.014023**). The uncertainty bands substantially overlap, so this release does **not** claim that either low-footprint tier is universally superior in fidelity.
87
 
88
+ **Positioning.** IQ1_M defines the current minimum-footprint frontier, while IQ2_S remains the released low-footprint balance point / lower-KLD alternative.
89
 
90
  > **Naming note:** `VeriLoop-E2-IQ1_M.gguf` is not a uniform 1-bit model. Its measured policy is **353 F32 + 1 IQ1_M + 2 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K = 851 tensors**, with **5.36 effective BPW**. `VeriLoop-E2-IQ2_S.gguf` is likewise a mixed-precision release rather than a uniform 2-bit model.
91
 
 
97
  | **Q8_0** | High-fidelity local | **28.596 GB / 26.632 GiB** | **46.86%** | **8.50 BPW** | Verified |
98
  | **Q6_K** | **Overall sweet spot** | **22.083 GB / 20.566 GiB** | **58.96%** | **6.57 BPW** | Verified / Release Quality PASS |
99
  | **Q5_K_M** | **Memory-quality sweet spot** | **20.364 GB / 18.965 GiB** | **62.16%** | **6.05 BPW** | Verified / Release Quality PASS |
100
+ | **Q4_K_M** | Intermediate mixed-precision candidate | **19.651 GB / 18.301 GiB** | **63.48%** | **5.84 BPW** | Quantitatively validated candidate |
101
  | **Q3_K_M** | Adjacent low-footprint alternative | **18.067 GB / 16.826 GiB** | **66.42%** | **5.37 BPW** | Full release PASS |
102
  | **IQ2_S** | **Released low-footprint sweet spot** | **18.037 GB / 16.799 GiB** | **66.48%** | **5.36 BPW** | **Full release PASS** |
103
+ | **IQ1_M** | **Ultra-low-footprint frontier candidate** | **18.028 GB / 16.790 GiB** | **66.50%** | **5.36 BPW** | **Quantitative fidelity PASS; local identity verified** |
104
 
105
  ## Quantization-retention benchmark
106
 
 
290
  hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-IQ2_S.gguf --local-dir .
291
  ```
292
 
293
+ > **IQ1_M is staged as the ultra-low-footprint frontier artifact.** Its quantitative evidence is PASS and its local SHA256 is frozen as `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`.
294
 
295
  Run IQ2_S:
296
 
 
317
  | Q4_K_M | **18.301 GiB** | Intermediate quantitative candidate |
318
  | Q3_K_M | **16.826 GiB** | Adjacent low-footprint release |
319
  | **IQ2_S** | **16.799 GiB** | **Released low-footprint sweet spot** |
320
+ | **IQ1_M** | **16.790 GiB** | **Ultra-low-footprint frontier candidate** |
321
 
322
  A **16.79 GiB model file does not imply full offload on a 16 GiB GPU**. Runtime memory also includes KV cache, compute buffers, allocator overhead, and optional MTP weights.
323
 
324
+ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128 41.999677 tok/s**), but there is no frozen paired BF16 throughput campaign. No universal speedup percentage is claimed for Q6/Q5/Q4/Q3/IQ2_S/IQ1_M.
325
 
326
  ## Release positioning
327
 
 
334
  | Q4_K_M | Intermediate mixed-precision candidate |
335
  | Q3_K_M | Adjacent low-footprint release |
336
  | **IQ2_S** | **Current released low-footprint sweet spot / lower-KLD frontier point** |
337
+ | **IQ1_M** | **Ultra-low-footprint frontier candidate** |
338
 
339
  ## Measurement boundaries
340
 
 
342
  - IQ1_M's **+0.3191% PPL** does not mean a 0.3191% drop on SWE-bench, Terminal-Bench, AIME, GPQA, or another downstream benchmark.
343
  - IQ2_S's **+0.3457% PPL** likewise does not imply a 0.3457% downstream-score loss.
344
  - IQ1_M and IQ2_S are both mixed-precision artifacts; the nominal tier name is not the model-wide effective bit width.
345
+ - IQ1_M has a frozen local SHA256 (`e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`) and quantitative PASS evidence.
346
  - The IQ1_M vs IQ2_S point-estimate differences are small relative to the reported uncertainty; this repository does not claim categorical fidelity superiority for either frontier point.
347
  - Native 262K context comes from the parent configuration; practical context depends on runtime memory.
348
  - MTP acceptance is prompt/workload-dependent.