malaiwah commited on
Commit
1e5c97a
·
verified ·
1 Parent(s): 6233eeb

card: explicit Hub lineage note (two sibling roots), cross-links to the measured family, discovery tags

Browse files
Files changed (1) hide show
  1. README.md +23 -0
README.md CHANGED
@@ -15,6 +15,10 @@ tags:
15
  - 8-bit
16
  - moe
17
  - reasoning
 
 
 
 
18
  ---
19
 
20
  # GLM-5.3-Flash-TR3-8bpw (K8)
@@ -102,6 +106,25 @@ combine order. Materialization receipt: bits 8, `complete`,
102
  `main_and_mtp_complete`, `nonrouted_native_exact`, 331,449,761,784 logical
103
  bytes, 37,152 routed choices, 1,618 native tensors.
104
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
105
  ## Credits
106
 
107
  Base model by [Z.ai](https://huggingface.co/zai-org). Quantization pipeline,
 
15
  - 8-bit
16
  - moe
17
  - reasoning
18
+ - text-generation
19
+ - fidelity
20
+ - kl-divergence
21
+ - exllamav3
22
  ---
23
 
24
  # GLM-5.3-Flash-TR3-8bpw (K8)
 
106
  `main_and_mtp_complete`, `nonrouted_native_exact`, 331,449,761,784 logical
107
  bytes, 37,152 routed choices, 1,618 native tensors.
108
 
109
+ ## Lineage on the Hub
110
+
111
+ Z.ai published two sibling roots for this model and neither declares the other:
112
+ [`zai-org/GLM-5.3-Flash`](https://huggingface.co/zai-org/GLM-5.3-Flash) (the
113
+ **FP8** release, where most traffic lands) and
114
+ [`zai-org/GLM-5.3-Flash-BF16`](https://huggingface.co/zai-org/GLM-5.3-Flash-BF16)
115
+ (the **BF16** weights). This quant declares BF16 as its `base_model` because
116
+ that is what it was actually quantized from — the FP8 release is a *sibling*
117
+ quantization of the same model, not our source, and it is the baseline we
118
+ measure against rather than build on. Quants that list FP8 as their base were
119
+ genuinely made from the FP8 weights; the trees differ for real reasons.
120
+
121
+ Related work on the same model, all measured on one panel in the
122
+ [quant-fidelity registry](https://huggingface.co/datasets/malaiwah/quant-fidelity-registry):
123
+ [brandonmusic 4bpw](https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw),
124
+ [0xSero Dione Q4](https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4),
125
+ [orcarouter MLX](https://huggingface.co/orcarouter/GLM-5.3-Flash-MLX).
126
+ Collection: [GLM-5.3-Flash — measured quants & fidelity](https://huggingface.co/collections/malaiwah/glm-53-flash-measured-quants-and-fidelity-6a91f253e7107818359f37c8).
127
+
128
  ## Credits
129
 
130
  Base model by [Z.ai](https://huggingface.co/zai-org). Quantization pipeline,