Myric commited on
Commit
82efefb
·
verified ·
1 Parent(s): bde4160

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +78 -0
README.md ADDED
@@ -0,0 +1,78 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: ibm-granite/granite-4.0-h-tiny
4
+ base_model_relation: quantized
5
+ pipeline_tag: text-generation
6
+ library_name: gguf
7
+ tags:
8
+ - gguf
9
+ - moe
10
+ - apex
11
+ - quantized
12
+ - granite
13
+ - mamba
14
+ - hybrid
15
+ - llama.cpp
16
+ ---
17
+
18
+ # Granite-4.0-H-Tiny — APEX GGUF
19
+
20
+ MoE-aware, mixed-precision **APEX** quantization of
21
+ [ibm-granite/granite-4.0-h-tiny](https://huggingface.co/ibm-granite/granite-4.0-h-tiny)
22
+ — IBM's **hybrid Mamba-2 / Transformer MoE** (`granitemoehybrid`): 40 layers = 36
23
+ Mamba-2 + 4 attention, a 64-routed + shared-expert MoE FFN on every layer, ~7B
24
+ total / ~1B active, Apache-2.0.
25
+
26
+ To my knowledge this is the first APEX quant of a Granite hybrid Mamba-2/MoE.
27
+ APEX assigns precision per tensor role and per layer; here that meant teaching the
28
+ recipe about the Mamba-2 mixer tensors the stock generator doesn't know (see *Method*).
29
+
30
+ ## Results
31
+
32
+ Perplexity on wikitext-2-raw (test, 200×512-token windows), `llama-perplexity`.
33
+
34
+ | File | Size | BPW | PPL | Δ vs bf16 |
35
+ |------|------|-----|-----|-----------|
36
+ | bf16 (reference) | 13 GB | 16.0 | 8.868 | — |
37
+ | **APEX i-quality** | **4.4 GB** | 5.40 | **8.901** | **+0.38%** |
38
+
39
+ Within **0.38% of full-precision perplexity at ~3× smaller**, and it runs comfortably
40
+ on modest hardware (~14 GB bf16 → 4.4 GB). Coherent, ~110 tok/s on a single GPU.
41
+ Built with a diverse imatrix (Bartowski `calibration_datav3`), full expert coverage.
42
+
43
+ ## Usage (llama.cpp)
44
+ ```bash
45
+ llama-cli -m granite-4.0-h-tiny-APEX-i-quality.gguf -ngl 999 -p "Hello"
46
+ llama-server -m granite-4.0-h-tiny-APEX-i-quality.gguf -ngl 999 --host 0.0.0.0 --port 8080
47
+ ```
48
+ Requires a llama.cpp build supporting the `granitemoehybrid` (a.k.a. `granitehybrid`)
49
+ architecture.
50
+
51
+ ## Method
52
+ APEX is a bit-allocation recipe over stock `llama-quantize --tensor-type-file`.
53
+ Granite-H needed the **Mamba-2 mixer** tensors added to the map, which the stock
54
+ APEX generator omits:
55
+ - **Mamba-2** (36 layers): `ssm_in`, `ssm_conv1d`, `ssm_out` at mixer precision
56
+ (Q6_K); 1-D state (`ssm_a`, `ssm_d`, `ssm_dt`, `ssm_norm`) left F32.
57
+ - **Attention** (4 layers): standard `attn_q/k/v/output` (Q6_K).
58
+ - **MoE** (all 40 layers): routed `ffn_*_exps` on a layer-depth precision gradient
59
+ (Q6_K/Q5_K/IQ4_XS), **shared experts `ffn_*_shexp` at Q8_0** (always active → protect),
60
+ router left high. Routed intermediate dim 512 is 256-divisible → no IQ4_NL workaround.
61
+
62
+ Config generation + patcher: `configs/`, `patch_granite_config.py`, `REPRODUCE.md`.
63
+ Baseline: IBM's own bf16 GGUF.
64
+
65
+ ### Note: pinning the Mamba-2 recurrence to Q8 doesn't help
66
+ A hand-roll variant pinning `ssm_in/out/conv1d` to Q8_0 scored **worse** at a
67
+ **larger** size (4.5 GB, PPL 8.913) — the SSM/recurrence tensors aren't precision-
68
+ sensitive here. Same null result we saw on Kimi-Linear's KDA. So **Q6_K is the right
69
+ choice**; don't spend bits protecting the linear-recurrence state. Not shipped.
70
+
71
+ ## Attribution & licenses
72
+ See [`LICENSE`](LICENSE) (Apache-2.0) and [`NOTICE`](NOTICE).
73
+ - Base: **IBM** ([@ibm-granite](https://huggingface.co/ibm-granite)) — [granite-4.0-h-tiny](https://huggingface.co/ibm-granite/granite-4.0-h-tiny) (Apache-2.0)
74
+ - Engine: **llama.cpp** ([@ggml-org](https://huggingface.co/ggml-org)) (MIT)
75
+ - APEX: **Ettore Di Giacinto / LocalAI** ([@mudler](https://huggingface.co/mudler)) — [localai-org/apex-quant](https://github.com/localai-org/apex-quant) (MIT)
76
+ - Calibration: **Bartowski** ([@bartowski](https://huggingface.co/bartowski)) — calibration_datav3
77
+
78
+ Unofficial community quantization; not affiliated with or endorsed by IBM.