File size: 6,117 Bytes
fe96519 18ddfa9 fe96519 18ddfa9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 | # Per-Tensor Quantization Arms — Methodology & Principles
> Companion technical document for [Confucius4-R2T2-Q4_K_M-GGUF](https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF).
> Details the general per-tensor quantization arms methodology, `--keep-type` semantics, cross-lingual validation protocols, and core architectural rules.
---
## 1. Why Per-Tensor Arms?
Standard model quantization tools apply uniform presets (e.g. standard Q4_K_M or Q5_K_M) across entire layers or architectures. For many speech models, however, uniform quantization either fails catastrophic quality checks or leaves substantial memory optimization on the table.
Building **quantization arms** means creating experimental variants that differ in exactly one block or layer type, isolating the quality, memory, and speed impact of each architectural component:
- Isolating which blocks represent non-negotiable precision floors.
- Finding which blocks can be aggressively quantized without measurable loss.
- Measuring real kernel execution speed across mixed-precision representations.
---
## 2. Converter Overrides & Mechanics
Per-tensor control requires specifying precision at the individual weight level:
- **Base Type:** Sets the baseline default type for all quantizable weights in the graph.
- **Explicit Overrides (`--keep-type <tensor>=<type>`):** Selectively overrides the precision of specific tensors.
- **Source Naming Contract:** Override rules must target the exact source parameter names (e.g. `thinker.model.layers.0.mlp.down_proj.weight`), not transformed runtime names.
- **Asserting Tensor Census:** Always assert a strict tensor census per architectural block before and after conversion to prevent silent fallbacks or unmatched patterns:
- Tower: 147 tensors
- Attention (`q`, `k`, `v`, `o_proj`): 112 tensors
- Gate / Up (`gate_proj`, `up_proj`): 56 tensors
- Down (`down_proj`): 28 tensors
- Embeddings: 1 tensor
- **Command-Line Limits:** When passing dozens of tensor overrides on Windows systems, drive conversion scripts from bash or JSON specifications rather than `cmd.exe` to avoid the 8191-character command-line length truncation limit.
---
## 3. Verification Protocol: Verify What Was Built
Never infer a model arm's composition from the conversion arguments alone—conversion flags can fail silently without throwing errors.
Always verify the resulting GGUF by inspecting the tensor header table:
1. Verify that all 28 `down_proj` weights are assigned their target type.
2. Confirm that `embed_tokens` was successfully overridden from the default F16 pin to the desired low-bit type.
3. Confirm that sensitive encoder layers retain their target mixed precision.
---
## 4. Cross-Lingual Validation Protocol
A critical lesson learned during quantization validation:
### English and Chinese Are Not a Screen
Speech models like R2T2 are trained with massive data allocations for English and Chinese, making them the model's most resilient input distributions.
- Multiple quantization configurations matched the Q8_0 reference baseline on English (`jfk`) and Chinese (`zh`), but completely broke down on non-target languages.
- For example, naive Q5_K matched Q8_0 on English and Chinese, but silently flipped to English when fed Russian audio.
- Other experimental configurations passed English and Chinese tests, but emitted English translations when given German audio.
### The Non-Target Probe Principle
Always evaluate quantization quality against languages with smaller training footprints (e.g. German, French, Russian). If a model switches languages or drops tokens on secondary languages, its representation space has degraded.
### Automated Detection
Rather than manually inspecting hours of transcripts, automated validation checks should monitor:
1. **Detected Language Output:** Grep the model runtime's detected language tag (`detected-language:`). A switch from `de` or `fr` to `en` immediately identifies representation collapse.
2. **Three-Tier Word Diffing:**
- **Raw:** Exact matching including punctuation and casing differences (benign).
- **Strict:** Lowercase, punctuation-stripped matching.
- **Loose:** Diacritics folded (NFD normalization). Flags true **content drift** (omitted, substituted, or hallucinated words).
3. **Empty Output Detection:** Low-bit matmuls can collapse abruptly to empty strings. Ensure output files are confirmed non-empty before running text diffs.
---
## 5. Architectural Principles That Generalize
The findings from the R2T2 quantization campaign reveal principles that apply broadly to modern transformer and speech architectures:
1. **Quantization Error is Cumulative, Not Per-Tensor:**
There is no fixed precision floor for a block in isolation. A block at Q4_K may work perfectly alone, but cause collapse when another block is also reduced. The total error budget is shared network-wide.
2. **Residual Stream Writes Compound Strictly:**
Blocks that write directly into the residual stream (`down_proj`) compound numerical error through every subsequent layer. Consequently, `down_proj` imposes the strictest precision requirement (Q6_K for R2T2).
3. **Non-Linearities Bound Error Propagation:**
Projections that feed into bounded activation functions (such as SwiGLU in `gate_proj` and `up_proj`) tolerate lower bitwidths (Q4_K) because the non-linearity bounds error growth.
4. **Lookup Tables Do Not Compound:**
Token embeddings (`embed_tokens`) are accessed via table lookup. Their quantization error is introduced once per token and does not propagate recurrently. Aggressively quantizing embeddings (even to Q2_K) saves substantial memory with virtually no degradation.
5. **Low-Bit Matmul Cliff:**
Matmul quantization does not degrade gracefully below Q4_K. Q3_K matmuls in R2T2 produced immediate empty output cliffs.
6. **Speed Does Not Track Bit Width Alone:**
Kernel optimization quality dominates hardware throughput. Q5_K performed ~20% slower than Q8_0 due to non-vectorized paths, whereas highly optimized Q4_K and Q6_K routines provided significant speedups.
|