Fidelity measurement: KL(reference || this quant) on frozen tokens, receipt-backed

#1
by malaiwah - opened

What this is. A third-party fidelity measurement of wrldsuksgo2mars/GLM-5.3-EXL3-K4-v1 @ 47af23347db743b4666d952e2eb48f2b01c3fede (exl3-trellis, 4 bits per weight as declared) against its unquantized reference, made with quant-fidelity-suite. Measured by malaiwah, not by the model's author. Every number below is read from a sealed receipt named at the bottom.

KL(reference ‖ candidate), mean tokenwise, nats 0.04480384821023634
top-1 agreement 0.939990229604299
KL median / p95 / p99 / max 0.0029634652346325514 / 0.19844349021462043 / 0.6973315018858828 / 9.842410402767543
reference root dataset malaiwah/glm53-fidelity-root-v1 @ 9c4a29ee10f393ed2fdbdb9262c1192ddb1507b4 (capture 9eba97dddb4ff2e2…)
panel panel--glm53.malaiwah.corpus5x5-v1, 25 contexts, 51175 scored positions
direction / vocabulary / accumulation KL(reference
method decode-and-run, weights only: exl3-trellis-decode-to-bf16 -- each exl3 payload group (trellis/suh/svh/codebook marker) decoded to bf16 per module on the capture device via exllamav3's transcribed codebooks (codebook per-module (read from each payload's own marker), declared 4 bits), non-routed tensors carried as shipped; same engine, schedule and device as the reference capture
determinism two fresh processes captured the candidate; both sealed captures carry content digest ba0e9beacbf0aaf3… (exact self-comparison 0.0)
comparability class strict

Scope (scope_digest): attn.o=quantized:fp8_e4m3@8|attn.other=native:bf16@16|attn.other=quantized:fp8_e4m3@8|attn.qkv=quantized:fp8_e4m3@8|embed_tokens=native:bf16@16|lm_head=native:bf16@16|mlp.down=quantized:fp8_e4m3@8|mlp.gate=quantized:fp8_e4m3@8|mlp.up=quantized:fp8_e4m3@8|moe.experts=quantized:exl3-mcg@4|moe.router=native:bf16@16|moe.shared_expert=quantized:fp8_e4m3@8|mtp=native:bf16@16|mtp=quantized:fp8_e4m3@8|norm=native:bf16@16|head=native|kv=bf16

Disclosures on the comparison receipt:

  • shared_reference_head (info): HEAD-1a: both captures declare the same lm_head tensor content digest (864f488a0074); one head applied to both sides is a shared APPLICATION of identical WEIGHTS.

What this number does not do. It is a same-lane distance from one reference capture on one panel. It does not rank this artifact against numbers measured on another panel, lane or reference, and a same-lane root does not retroactively upgrade rows measured against another teacher. Per-window scatter exceeds the gap between adjacent bit-widths; compare only within a group whose comparability keys match.

Receipts.

  • comparison receipt receipt_sha256 c10544b7f6f7e50eb0da9e6533b0090af9f9312dab35826f15a21b13875ee142
  • root-qualification receipt receipt_sha256 f08ce197a9e42e657f3451d75c7d066897ee80ef030620534ab1538e017488aa (canonical dataset_sha256 8d3b458e01a62c18578e037ca742b09943d7cefa079f5a6ae07225f859c6da14, repeat 2cd152919200f1c766f06a08bee6b4395a1d41492b36976792216a1ea61ec24c)
  • job job_id_full e3c643c5d8ba63290ff4a75391bbec4ba3d646b944261914140a73b2dd2cb0b5
  • candidate capture dataset published: malaiwah/glm53-fidelity-exl3-wrld-k4-v1 @ 9ef6de77ca2a534739ae314f498fa1019d74e235

Reproduce: fetch the two datasets named above and run fidelity-dataset compare --reference <root> --candidate <this> --force-compute; the receipt's estimator block is the exact recipe. Questions and corrections are welcome here; the registry files this as a third-party row (measured_by enumerated, never conflated with author-reported numbers).

Oh this is pretty cool… so this tool can be run locally by me for instance on my measly 2xRTX6000 even though those cards don’t have enough RAM run the BF16 original? Because the comparison dataset is a published artifact?

Would love to know how my GLM 5.3 Flash K3 does! And also some of my DS4 Flash experiments. I’ve been doing different mixing styles… will have a GLM flash 3:5:8 ratio 3.25bpw to test soon too.

Would be pretty interesting to have a site that actually catalogs the KL direct from hugging face for popular or submitted models…

Anyway thanks for the readout…

Yes, that is the goal. I am still improving on the tool, but I want to keep it so I can measure KLD of quants on my dev RTX5090 32GB (the tool does stream weights). Root datasets do require more vRAM but still not that much because of streaming. Measuring is kept at 1 card (no tensor parallelism) to have bit identical results whenever possible (TP influences order of operations).

The cli has a tool to retrieve the KLD numbers from the registry ( https://huggingface.co/datasets/malaiwah/quant-fidelity-registry ) . I am still working on getting this prettier on the HF side (like a dataset visualizer, because I get lost in the default one)

Follow-up, now that your K4 row and the rest of the GLM-5.3 group are in the public registry (malaiwah/quant-fidelity-registry, snapshot b58bca82). Two commands from a clone of https://github.com/malaiwah/quant-fidelity-suite, stock python3, no GPU, no token:

The whole GLM-5.3 comparability group (same root, same panel, same lane; ranks together):

$ bin/registry-view rows --panel panel--glm53.malaiwah.corpus5x5-v1 --registry hf
7 row(s) matched (filters hide groups; they never merge them)
COMPARABILITY GROUP cmp--fdbd312a2551db89
  mean_tokenwise_kld on panel--glm53.malaiwah.corpus5x5-v1 (same_stack, native_head)
  0.000000 nats GLM-5.3 BF16 (the official full-precision release)  [zai-org/GLM-5.3-BF16]
      id measurement--glm-5.3.bf16-selfcompare-floor.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class strict  measured_by self-measured:malaiwah
  0.022305 nats GLM-5.3 FP8 (the official block-scaled release)  [zai-org/GLM-5.3]
      id measurement--glm-5.3.fp8-dequantized.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah
  0.044804 nats wrldsuksgo2mars GLM-5.3 EXL3 K4 v1 (routed experts trellis K4, rest FP8)  [wrldsuksgo2mars/GLM-5.3-EXL3-K4-v1]
      id measurement--glm-5.3.exl3-k4-wrldsuksgo2mars.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah
  0.062842 nats davidsyoung GLM-5.3 EXL3 TR3 3.42bpw (routed experts trellis, TP4 rank-sharded)  [davidsyoung/GLM-5.3-EXL3-TR3-3.42bpw]
      id measurement--glm-5.3.exl3-tr3-3.42bpw-davidsyoung.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah
  0.073059 nats davidsyoung GLM-5.3 EXL3 TR3 3.25bpw (routed experts trellis, TP4 rank-sharded)  [davidsyoung/GLM-5.3-EXL3-TR3-3.25bpw]
      id measurement--glm-5.3.exl3-tr3-3.25bpw-davidsyoung.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah
  0.083833 nats davidsyoung GLM-5.3 EXL3 TR3 3.0bpw (routed experts trellis, TP4 rank-sharded)  [davidsyoung/GLM-5.3-EXL3-TR3-3.0bpw]
      id measurement--glm-5.3.exl3-tr3-3.0bpw-davidsyoung.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah
  0.102333 nats drowzeys keys-GLM-5.3-EXL3 (routed experts trellis 3.0 bpw, mcg/mul1)  [drowzeys/keys-GLM-5.3-EXL3]
      id measurement--glm-5.3.exl3-keys-drowzeys.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah
registry snapshot b58bca8293bb (fetched from the public HF dataset malaiwah/quant-fidelity-registry)

Asking the front gate about your repo (it refuses to re-spend on an artifact that is already measured at that exact revision):

$ bin/measure https://huggingface.co/wrldsuksgo2mars/GLM-5.3-EXL3-K4-v1 --plan-only
[1/9] target: wrldsuksgo2mars/GLM-5.3-EXL3-K4-v1 (live head)
[2/9] registry: registry snapshot b58bca8293bb (fetched from the public HF dataset malaiwah/quant-fidelity-registry)
[3/9] revision: 47af23347db7
  1 artifact record(s) for wrldsuksgo2mars/GLM-5.3-EXL3-K4-v1:
    [EXACT] artifact--wrldsuksgo2mars.glm-5.3-exl3-k4-v1
        measured at exactly this revision (47af23347d)
  1 published measurement row(s):
COMPARABILITY GROUP cmp--fdbd312a2551db89
  mean_tokenwise_kld on panel--glm53.malaiwah.corpus5x5-v1 (same_stack, native_head)
  (1 of 7 rows in this comparability group shown -- the registry holds 7; the whole group ranks together via bin/registry-view rows with no filters)
  0.044804 nats wrldsuksgo2mars GLM-5.3 EXL3 K4 v1 (routed experts trellis K4, rest FP8)  [wrldsuksgo2mars/GLM-5.3-EXL3-K4-v1]
      id measurement--glm-5.3.exl3-k4-wrldsuksgo2mars.corpus5x5-v1  panel panel--glm53.malaiwah.corpus5x5-v1  class advisory  measured_by self-measured:malaiwah

Each row links its sealed comparison receipt (the exact recipe: full vocabulary, fp64, KL(root || candidate), the head each side was replayed through) and the two published datasets it was computed from, so anyone can re-run fidelity-dataset compare and get the same tokenwise array.

The reference capture is a published dataset (malaiwah/glm53-fidelity-root-v1, 2.4 GB of hidden states + the head), so a quant is measured against it without ever loading the BF16 model; the candidate capture streams the quant one decoder layer at a time (peak 37.5 GB allocated for a GLM-5.3 candidate, 56.9 GB when the trellis decode runs on the device), which is inside one RTX PRO 6000.

FYI I made recent changes / corrections to the registry, but nothing that changes the score/measurement.

Your GLM-5.3-Flash K3 is measured: https://huggingface.co/wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1/discussions/1

Short version: KL(reference ‖ K3) = 0.0506 nats, top-1 0.931 (median 0.0039, p95 0.217, p99 0.718, max 6.99) on brandonmusic's 25-window final25 panel, 51,175 scored positions, against a fresh same-lane root capture of zai-org/GLM-5.3-Flash-BF16 (two cold H200 captures, bitwise identical, floor exactly 0.0 / top-1 1.0). Same engine, schedule, device and head on both sides, so the number is the quant's own distance with nothing subtracted. Both datasets are public (malaiwah/glm53-flash-fidelity-root-v1@bdd25fe0, malaiwah/glm53-flash-fidelity-exl3-wrld-k3-v1@e68c008c), and the comparison reproduces from them with fidelity-dataset compare --own-heads, no GPU.

Two things to keep straight when reading it beside older Flash numbers: this root starts a NEW comparability group on that panel (the 13 earlier Flash rows were scored against brandonmusic's stored teacher logits on a different lane and keep their own floor; nothing there is upgraded), and it is not comparable to the K4 row above either (different model, panel and root). It is filed as advisory for the same reason as the K4: the trellis payloads are decoded by this suite's transcription of exllamav3's codebooks, not by exllamav3's kernel.

Sign up or log in to comment