Myric commited on
Commit
44e9d5d
·
verified ·
1 Parent(s): 79ffe70

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +49 -0
README.md CHANGED
@@ -35,8 +35,15 @@ Perplexity on wikitext-2-raw (test, 200×512-token windows), `llama-perplexity`.
35
  |------|------|-----|-----|-----------|
36
  | bf16 (reference) | 13 GB | 16.0 | 8.868 | — |
37
  | **APEX i-quality** (ssm@Q6) | **4.4 GB** | 5.40 | **8.901** | **+0.38%** |
 
38
  | APEX hand-roll (ssm@Q8) | 4.5 GB | 5.54 | 8.913 | +0.51% |
39
 
 
 
 
 
 
 
40
  Within **0.38% of full-precision perplexity at ~3× smaller**, and it runs comfortably
41
  on modest hardware (~14 GB bf16 → 4.4 GB). Coherent, ~110 tok/s on a single GPU.
42
  Built with a diverse imatrix (Bartowski `calibration_datav3`), full expert coverage.
@@ -71,6 +78,48 @@ the linear-recurrence state helps: it doesn't — it's **larger and slightly wor
71
  hybrid architectures the SSM/recurrence tensors just aren't precision-sensitive. The
72
  hand-roll is included to document the experiment; `i-quality` is the recommended tier.
73
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
74
  ## Attribution & licenses
75
  See [`LICENSE`](LICENSE) (Apache-2.0) and [`NOTICE`](NOTICE).
76
  - Base: **IBM** ([@ibm-granite](https://huggingface.co/ibm-granite)) — [granite-4.0-h-tiny](https://huggingface.co/ibm-granite/granite-4.0-h-tiny) (Apache-2.0)
 
35
  |------|------|-----|-----|-----------|
36
  | bf16 (reference) | 13 GB | 16.0 | 8.868 | — |
37
  | **APEX i-quality** (ssm@Q6) | **4.4 GB** | 5.40 | **8.901** | **+0.38%** |
38
+ | **APEX i-quality — torch-imatrix** | **4.4 GB** | 5.40 | **8.864** | **−0.04%** |
39
  | APEX hand-roll (ssm@Q8) | 4.5 GB | 5.54 | 8.913 | +0.51% |
40
 
41
+ The **torch-imatrix** row is the *same* APEX i-quality recipe, quantized with an
42
+ importance matrix computed independently in PyTorch (see [*The torch-imatrix
43
+ variant*](#the-torch-imatrix-variant)) instead of `llama-imatrix`. Its PPL (8.864)
44
+ is inside the ±0.11 error bar of both bf16 and the standard i-quality — i.e. the
45
+ two imatrices are interchangeable in quality.
46
+
47
  Within **0.38% of full-precision perplexity at ~3× smaller**, and it runs comfortably
48
  on modest hardware (~14 GB bf16 → 4.4 GB). Coherent, ~110 tok/s on a single GPU.
49
  Built with a diverse imatrix (Bartowski `calibration_datav3`), full expert coverage.
 
78
  hybrid architectures the SSM/recurrence tensors just aren't precision-sensitive. The
79
  hand-roll is included to document the experiment; `i-quality` is the recommended tier.
80
 
81
+ ## The torch-imatrix variant
82
+
83
+ `granite-4.0-h-tiny-APEX-i-quality-torch.gguf` (+ `granite-4.0-h-tiny-torch.imatrix`)
84
+ is a companion build that answers one question: **can the calibration imatrix be
85
+ produced without llama.cpp, and does that change the result?**
86
+
87
+ **What it is.** An importance matrix (imatrix) records, per weight tensor, the
88
+ per-input-channel sum of squared activations over calibration text — it tells
89
+ `llama-quantize` where to spend bits. Normally you get it from `llama-imatrix`,
90
+ which needs the whole model resident in RAM+VRAM. Here it was instead computed
91
+ **directly from the Hugging Face model in PyTorch**, by registering forward hooks
92
+ on every matmul (attention, Mamba-2 in/out projections, router, shared experts,
93
+ and per-expert routed FFNs) and accumulating Σx² as calibration text streams
94
+ through. The generator is **band-serialized**: it loads a few decoder layers at a
95
+ time, runs all calibration chunks through them, caches activations, frees them,
96
+ and moves on — so peak VRAM is a few layers, **not the whole model** (this run
97
+ peaked at ~6.4 GB on a 16 GB card, vs the ~14 GB a full-resident forward needs).
98
+ That decouples imatrix generation from llama.cpp's memory model: you can calibrate
99
+ a model far larger than your GPU by streaming it in bands.
100
+
101
+ **Everything else is identical.** Same bf16 base, same APEX `--tensor-type-file`
102
+ recipe, same `calibration_datav3`, same 126×512 calibration windows. The *only*
103
+ difference from the standard `i-quality` file is which tool produced the imatrix.
104
+ The two quants therefore differ **only** in the tensors whose quantization
105
+ consumes an imatrix (the K-/I-quant expert, attention and Mamba weights); the
106
+ Q8_0 shared experts and F32 tensors are bit-identical.
107
+
108
+ **Validation.** The PyTorch imatrix matches `llama-imatrix`'s output tensor-for-
109
+ tensor (368/368 names) with **median per-tensor correlation 0.995**. The handful
110
+ of lower-correlation tensors are the post-nonlinearity inputs (`ffn_down_shexp`,
111
+ `attn_output`) where transformers' naive Mamba scan and SiLU-gated intermediates
112
+ differ numerically from llama.cpp's kernels — a known wrinkle that does not affect
113
+ quality, since the quantizer only needs *relative* per-channel importance. The
114
+ proof is in the table above: the torch-imatrix quant scores **PPL 8.864**, inside
115
+ the ±0.11 error bar of both bf16 and the standard i-quality. The two imatrices are
116
+ interchangeable.
117
+
118
+ **Pick whichever you like** — they are equivalent in quality. The standard
119
+ `i-quality` is the reference; the `torch` build is provided for anyone who wants a
120
+ llama.cpp-free, VRAM-bounded path to the same result (e.g. calibrating very large
121
+ MoEs on modest hardware).
122
+
123
  ## Attribution & licenses
124
  See [`LICENSE`](LICENSE) (Apache-2.0) and [`NOTICE`](NOTICE).
125
  - Base: **IBM** ([@ibm-granite](https://huggingface.co/ibm-granite)) — [granite-4.0-h-tiny](https://huggingface.co/ibm-granite/granite-4.0-h-tiny) (Apache-2.0)