# VeriLoop E2 — Q8_0 Quality Card **Release type:** high-fidelity Q8_0 GGUF **Baseline:** canonical VeriLoop E2 BF16 GGUF **Purpose:** quantify BF16 → Q8_0 distortion under a fixed llama.cpp protocol **Formal parent-model benchmark re-run:** no **Release status:** **PASS** **Positioning:** **High-fidelity / near-reference Q8_0 tier**. This card evaluates quantization fidelity directly against BF16; it does not re-label the parent model's downstream benchmark scores as Q8-specific results. --- ## 1. Artifact Identity | Property | Q8_0 main | Optional MTP draft | |---|---|---| | Filename | `VeriLoop-E2-Q8_0.gguf` | `mtp-VeriLoop-E2-Q8_0.gguf` | | Exact bytes | 28595765600 | 3287070240 | | Binary size | 26.632 GiB | 3.061 GiB | | Decimal size | 28.596 GB | 3.287 GB | | SHA256 | `see `Q8_RELEASE_MANIFEST.json`` | `see `Q8_RELEASE_MANIFEST.json`` | | Status | **Verified** | **Verified** | BF16 reference SHA256: ```text 11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d ``` Release finalized: ```text see `Q8_RELEASE_MANIFEST.json` ``` --- ## 2. Quantization Build | Item | Value | |---|---| | Parent model | VeriLoop E2 | | Base family | Qwen3.8-27B | | HF architecture | `Qwen3_5ForConditionalGeneration` | | Quant type | `Q8_0` | | Main tensors | 851 | | Q8_0 tensors | 498 | | Retained F32 tensors | 353 | | Source model size reported by quantizer | 51,305.09 MiB | | Quantized model size reported by quantizer | 27,260.56 MiB | | Effective density reported by quantizer | **8.50 BPW** | | Size reduction vs BF16 GGUF | **46.86%** | | Quantization threads | 25 | | Quantization time | **38.315 s** | | Quantization return code | **0** | | Main build gate | **PASS** | | llama.cpp quantization/QC revision | `42916d83f4a225e56709f873aa8050ac11f5b6a4` | The 353 F32 tensors are intentionally retained by the llama.cpp quantized layout; the release does not claim every tensor is stored as Q8_0. --- ## 3. Fixed BF16 → Q8_0 Comparison Protocol | Item | Value | |---|---| | Reference model | VeriLoop E2 BF16 GGUF | | Candidate | VeriLoop E2 Q8_0 GGUF | | Corpus | WikiText-2 raw test | | Context length | 2,048 | | Chunks | 8 | | Seed | 42 | | GPU layers | 40 | | Flash Attention | Off | | KV cache | F16 / F16 | | Batch size | 512 | | Micro-batch size | 512 | | Evaluation executable | `llama-perplexity` | | Reference-logit method | BF16 logits persisted with `--kl-divergence-base` | | Candidate comparison | Q8_0 evaluated with `--kl-divergence` against the frozen BF16 logit file | This protocol is a **paired quantization-fidelity test**. It is not a new leaderboard campaign. --- ## 4. Quantization Fidelity ### Perplexity | Metric | Value | |---|---:| | BF16 mean PPL | **4.840423 ± 0.119931** | | Q8_0 mean PPL | **4.843536 ± 0.120062** | | Absolute ΔPPL | **+0.003113 ± 0.003651** | | PPL ratio | **1.000643 ± 0.000754** | | Relative PPL increase | **+0.0643%** | | Cor(ln PPL(Q8), ln PPL(BF16)) | **99.95%** | ### KL divergence | Metric | Value | |---|---:| | Mean KLD | **0.002176 ± 0.000668** | | Median KLD | **0.000276** | | 90th percentile | **0.001548** | | 95th percentile | **0.002856** | | 99th percentile | **0.009561** | | 99.9th percentile | **0.201704** | | Maximum | **4.066900** | ### Token-probability stability | Metric | Value | |---|---:| | Mean Δp | **−0.007 ± 0.014%** | | RMS Δp | **1.251 ± 0.137%** | | Same top-p | **98.815 ± 0.120%** | ### Interpretation The measured Q8_0 distortion is small under the fixed protocol: - predictive loss rises by **0.0643% relative to BF16**; - average distributional divergence is **0.002176 KLD**; - the top-probability token remains the same **98.815%** of the time; - mean probability shift is close to zero, while RMS Δp quantifies the remaining quantization noise. **Important:** the +0.0643% figure is a relative perplexity increase, not a universal downstream benchmark or capability-loss percentage. --- ## 5. Release Gates | Gate | Predeclared release threshold | Measured | Result | |---|---:|---:|---| | PPL ratio | ≤ 1.005 | **1.000643** | **PASS** | | Mean KLD | ≤ 0.005 | **0.002176** | **PASS** | | Same top-p | ≥ 97.0% | **98.815%** | **PASS** | | Main tensor count | 851 | **851** | **PASS** | | Non-zero / structural audit | required | **PASS** | **PASS** | | llama.cpp model load | required | **PASS** | **PASS** | | OpenAI-compatible generation | required | **PASS** | **PASS** | | Optional MTP loader compatibility | required for MTP artifact | **PASS** | **PASS** | | Optional MTP speculative draft smoke | required for MTP artifact | **PASS** | **PASS** | These thresholds are VeriLoop internal release-QA criteria and are not presented as universal quantization standards. --- ## 6. Runtime Compatibility Validated llama.cpp revision: ```text 42916d83f4a225e56709f873aa8050ac11f5b6a4 ``` Validated MTP compatibility smoke: | Item | Value | |---|---| | Main model | `VeriLoop-E2-Q8_0.gguf` | | Draft model | `mtp-VeriLoop-E2-Q8_0.gguf` | | Context | 8,192 | | Main GPU offload | enabled | | Draft GPU layers | 0 | | Model load | **PASS** | | Generation | **PASS** | | Draft tokens proposed | 3 | | Draft tokens accepted | 3 | An earlier 32K all-GPU main+draft attempt exhausted available CUDA memory while allocating the draft runtime buffer. It did **not** produce a tensor-layout or GGUF-loader error. The resource-conservative 8K main-GPU/draft-CPU configuration loaded and generated successfully. This card therefore distinguishes **format/runtime compatibility** from **hardware-specific memory fit**. --- ## 7. Size Efficiency | Artifact | Size | |---|---:| | BF16 reference GGUF | **50.113 GiB** | | Q8_0 GGUF | **26.632 GiB** | | Reduction | **46.86%** | The final exact Q8_0 byte count is 28595765600 bytes. --- ## 8. Throughput Note A standalone Q8_0 `llama-bench` run produced: | Workload | Q8_0 | |---|---:| | Prompt processing (`pp512`) | **3102.722776 tok/s** | | Token generation (`tg128`) | **41.999677 tok/s** | The paired BF16 speed run did not complete under the original full-offload benchmark configuration, so **no BF16→Q8 speedup factor is claimed**. These Q8_0 throughput numbers are descriptive only and are not part of the quality gate. --- ## 9. Parent-Model Scores The public VeriLoop E2 benchmark record belongs to the parent release. Q8_0 was not re-scored across the full benchmark suite for this quantization card. That separation is intentional: ```text Parent-model evaluation → establishes model capability BF16 vs Q8_0 paired PPL/KLD/token analysis → establishes quantization fidelity ``` This prevents format conversion and model capability from being conflated. --- ## 10. Final Decision ```text MAIN_Q8_BUILD=PASS Q8_STRUCTURE_GATE=PASS Q8_QUANTITATIVE_QUALITY_GATE=PASS Q8_STOCK_LLAMA_CPP_COMPATIBILITY=PASS MTP_RUNTIME_GATE=PASS Q8_RELEASE_QUALITY_GATE=PASS ``` **Release conclusion:** Q8_0 is accepted as the high-fidelity local-deployment quantization of VeriLoop E2 under the frozen release protocol. The strongest numerical statement supported by the evidence is: > **Compared with the canonical BF16 GGUF under the identical fixed protocol, Q8_0 reduces file size by 46.86%, increases mean perplexity by 0.0643%, yields mean KLD 0.002176, and preserves the same top-probability token 98.815% of the time.** No broader capability-loss percentage is inferred from these statistics.