VeriLoop-E2-GGUF / QUALITY_CARD.md
ConorWang's picture
Rename Q8_QUALITY_CARD.md to QUALITY_CARD.md
c874f52 verified
|
Raw History Blame Contribute Delete
7.46 kB

VeriLoop E2 — Q8_0 Quality Card

Release type: high-fidelity Q8_0 GGUF
Baseline: canonical VeriLoop E2 BF16 GGUF
Purpose: quantify BF16 → Q8_0 distortion under a fixed llama.cpp protocol
Formal parent-model benchmark re-run: no
Release status: PASS

Positioning: High-fidelity / near-reference Q8_0 tier. This card evaluates quantization fidelity directly against BF16; it does not re-label the parent model's downstream benchmark scores as Q8-specific results.


1. Artifact Identity

Property Q8_0 main Optional MTP draft
Filename VeriLoop-E2-Q8_0.gguf mtp-VeriLoop-E2-Q8_0.gguf
Exact bytes 28595765600 3287070240
Binary size 26.632 GiB 3.061 GiB
Decimal size 28.596 GB 3.287 GB
SHA256 see Q8_RELEASE_MANIFEST.json`` see Q8_RELEASE_MANIFEST.json``
Status Verified Verified

BF16 reference SHA256:

11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d

Release finalized:

see `Q8_RELEASE_MANIFEST.json`

2. Quantization Build

Item Value
Parent model VeriLoop E2
Base family Qwen3.8-27B
HF architecture Qwen3_5ForConditionalGeneration
Quant type Q8_0
Main tensors 851
Q8_0 tensors 498
Retained F32 tensors 353
Source model size reported by quantizer 51,305.09 MiB
Quantized model size reported by quantizer 27,260.56 MiB
Effective density reported by quantizer 8.50 BPW
Size reduction vs BF16 GGUF 46.86%
Quantization threads 25
Quantization time 38.315 s
Quantization return code 0
Main build gate PASS
llama.cpp quantization/QC revision 42916d83f4a225e56709f873aa8050ac11f5b6a4

The 353 F32 tensors are intentionally retained by the llama.cpp quantized layout; the release does not claim every tensor is stored as Q8_0.


3. Fixed BF16 → Q8_0 Comparison Protocol

Item Value
Reference model VeriLoop E2 BF16 GGUF
Candidate VeriLoop E2 Q8_0 GGUF
Corpus WikiText-2 raw test
Context length 2,048
Chunks 8
Seed 42
GPU layers 40
Flash Attention Off
KV cache F16 / F16
Batch size 512
Micro-batch size 512
Evaluation executable llama-perplexity
Reference-logit method BF16 logits persisted with --kl-divergence-base
Candidate comparison Q8_0 evaluated with --kl-divergence against the frozen BF16 logit file

This protocol is a paired quantization-fidelity test. It is not a new leaderboard campaign.


4. Quantization Fidelity

Perplexity

Metric Value
BF16 mean PPL 4.840423 ± 0.119931
Q8_0 mean PPL 4.843536 ± 0.120062
Absolute ΔPPL +0.003113 ± 0.003651
PPL ratio 1.000643 ± 0.000754
Relative PPL increase +0.0643%
Cor(ln PPL(Q8), ln PPL(BF16)) 99.95%

KL divergence

Metric Value
Mean KLD 0.002176 ± 0.000668
Median KLD 0.000276
90th percentile 0.001548
95th percentile 0.002856
99th percentile 0.009561
99.9th percentile 0.201704
Maximum 4.066900

Token-probability stability

Metric Value
Mean Δp −0.007 ± 0.014%
RMS Δp 1.251 ± 0.137%
Same top-p 98.815 ± 0.120%

Interpretation

The measured Q8_0 distortion is small under the fixed protocol:

  • predictive loss rises by 0.0643% relative to BF16;
  • average distributional divergence is 0.002176 KLD;
  • the top-probability token remains the same 98.815% of the time;
  • mean probability shift is close to zero, while RMS Δp quantifies the remaining quantization noise.

Important: the +0.0643% figure is a relative perplexity increase, not a universal downstream benchmark or capability-loss percentage.


5. Release Gates

Gate Predeclared release threshold Measured Result
PPL ratio ≤ 1.005 1.000643 PASS
Mean KLD ≤ 0.005 0.002176 PASS
Same top-p ≥ 97.0% 98.815% PASS
Main tensor count 851 851 PASS
Non-zero / structural audit required PASS PASS
llama.cpp model load required PASS PASS
OpenAI-compatible generation required PASS PASS
Optional MTP loader compatibility required for MTP artifact PASS PASS
Optional MTP speculative draft smoke required for MTP artifact PASS PASS

These thresholds are VeriLoop internal release-QA criteria and are not presented as universal quantization standards.


6. Runtime Compatibility

Validated llama.cpp revision:

42916d83f4a225e56709f873aa8050ac11f5b6a4

Validated MTP compatibility smoke:

Item Value
Main model VeriLoop-E2-Q8_0.gguf
Draft model mtp-VeriLoop-E2-Q8_0.gguf
Context 8,192
Main GPU offload enabled
Draft GPU layers 0
Model load PASS
Generation PASS
Draft tokens proposed 3
Draft tokens accepted 3

An earlier 32K all-GPU main+draft attempt exhausted available CUDA memory while allocating the draft runtime buffer. It did not produce a tensor-layout or GGUF-loader error. The resource-conservative 8K main-GPU/draft-CPU configuration loaded and generated successfully.

This card therefore distinguishes format/runtime compatibility from hardware-specific memory fit.


7. Size Efficiency

Artifact Size
BF16 reference GGUF 50.113 GiB
Q8_0 GGUF 26.632 GiB
Reduction 46.86%

The final exact Q8_0 byte count is 28595765600 bytes.


8. Throughput Note

A standalone Q8_0 llama-bench run produced:

Workload Q8_0
Prompt processing (pp512) 3102.722776 tok/s
Token generation (tg128) 41.999677 tok/s

The paired BF16 speed run did not complete under the original full-offload benchmark configuration, so no BF16→Q8 speedup factor is claimed. These Q8_0 throughput numbers are descriptive only and are not part of the quality gate.


9. Parent-Model Scores

The public VeriLoop E2 benchmark record belongs to the parent release. Q8_0 was not re-scored across the full benchmark suite for this quantization card.

That separation is intentional:

Parent-model evaluation
    → establishes model capability

BF16 vs Q8_0 paired PPL/KLD/token analysis
    → establishes quantization fidelity

This prevents format conversion and model capability from being conflated.


10. Final Decision

MAIN_Q8_BUILD=PASS
Q8_STRUCTURE_GATE=PASS
Q8_QUANTITATIVE_QUALITY_GATE=PASS
Q8_STOCK_LLAMA_CPP_COMPATIBILITY=PASS
MTP_RUNTIME_GATE=PASS
Q8_RELEASE_QUALITY_GATE=PASS

Release conclusion: Q8_0 is accepted as the high-fidelity local-deployment quantization of VeriLoop E2 under the frozen release protocol.

The strongest numerical statement supported by the evidence is:

Compared with the canonical BF16 GGUF under the identical fixed protocol, Q8_0 reduces file size by 46.86%, increases mean perplexity by 0.0643%, yields mean KLD 0.002176, and preserves the same top-probability token 98.815% of the time.

No broader capability-loss percentage is inferred from these statistics.