File size: 10,991 Bytes
eca8325
ec1bd06
eca8325
 
bec001f
 
 
 
 
 
 
 
ec1bd06
 
eca8325
 
50287a7
 
 
 
578c75d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eca8325
 
ec1bd06
 
 
 
 
eca8325
bec001f
 
 
 
 
 
 
 
eca8325
 
ec1bd06
 
 
 
 
 
eca8325
 
 
578c75d
ec1bd06
 
 
eca8325
ec1bd06
 
 
 
 
 
 
 
 
 
 
 
 
 
eca8325
ec1bd06
 
 
 
 
eca8325
 
ec1bd06
 
 
 
 
eca8325
ec1bd06
 
 
 
eca8325
 
ec1bd06
 
 
 
 
 
 
 
 
 
 
 
 
eca8325
 
 
 
 
ec1bd06
 
 
 
 
 
 
 
 
 
 
 
eca8325
caaf9e6
eca8325
caaf9e6
 
 
 
 
 
ec1bd06
 
 
 
 
 
 
 
 
 
 
 
 
 
eca8325
ec1bd06
 
 
 
 
eca8325
 
 
 
ec1bd06
 
 
 
 
90f7b21
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1e69c6f
90f7b21
 
 
cf39760
 
 
90f7b21
cf39760
 
 
 
 
 
 
caaf9e6
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
---
license: apache-2.0
base_model: google/gemma-4-31B-it
tags:
  - gguf
  - rotorquant
  - kv-cache-quantization
  - gemma
  - gemma4
  - instruct
  - llama-cpp
  - quantized
library_name: gguf
pipeline_tag: image-text-to-text
---

> [!WARNING]
> **Fork compatibility (2026-07-07):** the `llama-cpp-turboquant` fork is currently based on a llama.cpp revision that **predates `gemma4` architecture support** β€” it fails with `unknown model architecture: 'gemma4'` and cannot run this model at all. Until the fork rebases, use **mainline llama.cpp** (which loads this GGUF fine with standard KV-cache types); the RotorQuant/TurboQuant KV-cache options are not usable with gemma-4 yet.
<!-- gemma4-fork-note -->

> [!TIP]
> **KV-cache quantization without any fork (recommended, 2026):** upstream
> llama.cpp/Ollama now cover this natively β€” use `-ctk q8_0 -ctv q8_0`
> (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
> `-ctk q4_0 -ctv q4_0` (~quarter memory, β‰ˆ7.6% perplexity increase). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
> K and V types symmetric to stay on the fast fused Flash-Attention path.
> Since April 2026, mainline llama.cpp also applies Hadamard rotation to
> KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
> which greatly improves low-bit KV quality (opt-out:
> `LLAMA_ATTN_ROT_DISABLE=1`).
>
> The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
> TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
> is unmaintained relative to mainline. It is NOT required to use this model.
<!-- kv-upstream-note -->

# gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M

GGUF Q4_K_M weight-quantized variant of [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) optimised for use with **RotorQuant** KV cache compression via a dedicated llama.cpp fork.

> **Important:** RotorQuant KV cache types (`planar3`, `iso3`) are **not** available in upstream llama.cpp, standard Ollama, or LM Studio.
> They require a [specific llama.cpp fork](https://github.com/johndpope/llama-cpp-turboquant/tree/feature/planarquant-kv-cache).
> The GGUF file itself is a standard GGUF and works with any llama.cpp-compatible runtime using normal KV cache types (f16, q8_0, q4_0, etc.).

## Hardware compatibility

| Device | VRAM / RAM | Recommendation |
| --- | --- | --- |
| CPU host with β‰₯18 GB RAM | ~18.4 GB | works via llama.cpp; slower than GPU but no accelerator required |
| Apple Silicon (Metal) | ~20.0 GB | llama.cpp Metal backend; fast on M-series unified memory |
| NVIDIA GPU (partial offload) | split between GPU + RAM | offload as many layers as VRAM allows; rest on CPU |

## Overview

This model combines two independent compression techniques:

| Technique | What it does | Requirement |
|-----------|-------------|-------------|
| **GGUF Q4_K_M weight quantization** | Reduces model size from ~62 GB (BF16) to ~16.7 GB | Any llama.cpp-compatible runtime |
| **RotorQuant KV cache compression** β€” block-diagonal Clifford-algebra rotors for 3-bit KV cache (`--cache-type-k iso3 --cache-type-v iso3`) | Block-diagonal rotations / random rotation for compressed KV cache | [llama-cpp-turboquant fork](https://github.com/johndpope/llama-cpp-turboquant/tree/feature/planarquant-kv-cache) only |

## Quickstart

### Option A β€” RotorQuant KV cache (experimental fork β€” not required)

You must build from the RotorQuant-enabled llama.cpp fork:

```bash
# Clone and build the fork
git clone https://github.com/johndpope/llama-cpp-turboquant.git
cd llama-cpp-turboquant && git checkout feature/planarquant-kv-cache

# CUDA (Windows/Linux)
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

# Metal (Apple Silicon)
cmake -B build -DGGML_METAL=ON -DGGML_METAL_EMBED_LIBRARY=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

# Run with RotorQuant KV cache
./build/bin/llama-cli -m gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M.gguf \
  --cache-type-k iso3 --cache-type-v iso3 \
  -ngl 99 -fa \
  -p "Explain quantum computing"

# Or run as a server
./build/bin/llama-server -m gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M.gguf \
  --cache-type-k iso3 --cache-type-v iso3 \
  -ngl 99 -fa --jinja
```

### Option B β€” With standard llama.cpp / LM Studio / Ollama

The GGUF works as a normal quantised model. You won't get RotorQuant-specific KV cache benefits, but standard KV cache quantization (q8_0, q4_0) still reduces VRAM significantly.

**llama.cpp (upstream)**
```bash
llama-cli -m gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M.gguf \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  -ngl 99 -fa \
  -p "Explain quantum computing"
```

**LM Studio**
1. Download the GGUF file and load in LM Studio.
2. Enable **Developer Mode** (Settings β†’ Developer).
3. In the model loader's advanced settings, set **Flash Attention** to ON.
4. Set **K Cache Quantization** and **V Cache Quantization** to `q8_0` (or `q4_0` for more aggressive VRAM savings).
5. Note: LM Studio does not currently support RotorQuant's `iso3` cache types. Track [this feature request](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1719) for updates.

**Ollama**
```bash
# Standard Ollama does not support RotorQuant cache types.
# Use with default or q8_0 KV cache via OLLAMA_KV_CACHE_TYPE=q8_0
OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_FLASH_ATTENTION=1 ollama run majentik/gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M
```

## Specifications

| Property | Value |
|----------|-------|
| Base Model | [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) |
| Architecture | Dense transformer, instruct-tuned |
| Parameters | 31B (all active, dense) |
| Context Length | 128K |
| Weight Quantization | GGUF Q4_K_M (popular 4-bit, best quality/size tradeoff) |
| Original Size (BF16) | ~62 GB |
| Quantized File Size | ~16.7 GB |
| KV Cache (RotorQuant) | 3-bit via `--cache-type-k iso3 --cache-type-v iso3` (fork only) |
| KV Cache (standard) | q8_0, q4_0, f16, etc. (any llama.cpp runtime) |
| License | apache-2.0 |
| Modalities | Text + Image (image-text-to-text) |
| Compatible Runtimes | llama.cpp, LM Studio, Ollama, koboldcpp |

## About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's **release labels**, not distinct
quantization algorithms β€” for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured. The KV-cache fork these
labels originally referred to is legacy; for KV-cache memory savings use the
upstream options described above (`-ctk/-ctv q8_0`, `OLLAMA_KV_CACHE_TYPE`).

## Current Status of RotorQuant in the Ecosystem

| Runtime | RotorQuant Support | Standard KV Quant |
|---------|---------------------|-------------------|
| llama.cpp (upstream) | ❌ Not merged | βœ… q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1 |
| llama-cpp-turboquant fork | βœ… planar3, iso3 | βœ… All standard types |
| LM Studio | ❌ [Requested](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1719) | βœ… Via advanced settings |
| Ollama | ❌ Not supported | βœ… Via OLLAMA_KV_CACHE_TYPE |
| koboldcpp | ❌ Not supported | βœ… Standard types |

## Recommended Settings

For VRAM-constrained setups, standard q8_0 KV cache quantization already halves KV cache memory with negligible quality impact. Flash Attention should always be enabled β€” it is required for V cache quantization and improves memory efficiency regardless.

| VRAM | Suggested Configuration |
|------|------------------------|
| 24 GB (RTX 4090) | Q4_K_M + q8_0 KV cache + Flash Attention, 8K–16K context |
| 16 GB | Q4_K_M + q4_0 KV cache + Flash Attention, 4K–8K context |
| 48+ GB | Q4_K_M + f16 KV cache, full 32K+ context |

## See Also

- [RotorQuant GitHub](https://github.com/scrya-com/rotorquant)
- [llama-cpp-turboquant fork](https://github.com/johndpope/llama-cpp-turboquant/tree/feature/planarquant-kv-cache)
- [TurboQuant llama.cpp discussion](https://github.com/ggml-org/llama.cpp/discussions/20969)
- [TurboQuant paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
- [Base model: google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it)
- [gemma-4-31B-it announcement](https://blog.google/technology/developers/gemma-4/)

## Quant trade-off (GGUF lane)

| Quant | Approx size | Use case | Recommendation |
|---|---|---|---|
| Q2_K | ~17 GB | Lossy, low-RAM CPU/edge | Resource-constrained inference |
| Q3_K_M | ~19 GB | Smaller-than-Q4, modest quality drop | Edge devices with ~16 GB RAM |
| IQ4_XS | ~16 GB | Importance-quant 4-bit, smaller than Q4_K_M | Best size/quality at 4-bit |
| **Q4_K_M** | ~23 GB | Balanced default | **Recommended for most users** |
| Q5_K_M | ~24 GB | Higher fidelity than Q4 | Quality-sensitive applications |
| Q6_K | ~28 GB | Approaching FP16 quality | High-fidelity CPU/edge |
| Q8_0 | ~32 GB | Near-lossless reference | Fidelity-critical work |
| MXFP4_MOE | ~17 GB | Microscaling FP4 (MoE-aware) | vLLM / transformers users |

(Current variant β€” **Q4_K_M** β€” is bolded.)

## Variants in this family

(Showing 14 sibling variants under `majentik/gemma-4-31b-it-*`. The current variant β€” `RotorQuant-GGUF-Q4_K_M` β€” is **bolded**.)

| Variant | Runtime | Approx size | Use case |
|---|---|---|---|
| [RotorQuant-GGUF-IQ4_XS](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-gguf-IQ4_XS) | llama.cpp | ~27 GB | Lossy 4-bit, low-RAM CPU/edge |
| [RotorQuant-GGUF-Q2_K](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-gguf-Q2_K) | llama.cpp | ~19 GB | Lossy, low-RAM CPU/edge |
| [RotorQuant-GGUF-Q3_K_M](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-gguf-Q3_K_M) | llama.cpp | ~24 GB | Smaller 3-bit, CPU-friendly |
| **RotorQuant-GGUF-Q4_K_M** | llama.cpp | ~34 GB | Balanced default |
| [RotorQuant-GGUF-Q5_K_M](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-gguf-Q5_K_M) | llama.cpp | ~41 GB | Higher fidelity, more RAM |
| [RotorQuant-GGUF-Q8_0](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-gguf-Q8_0) | llama.cpp | ~65 GB | Near-lossless reference |
| [RotorQuant-MLX-2bit](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-mlx-2bit) | mlx-lm | ~9.9 GB | Apple Silicon, smallest |
| [RotorQuant-MLX-4bit](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-mlx-4bit) | mlx-lm | ~19 GB | Apple Silicon balanced |
| [RotorQuant-MLX-8bit](https://huggingface.co/majentik/gemma-4-31b-it-rotorquant-mlx-8bit) | mlx-lm | ~37 GB | Apple Silicon reference |
| [TurboQuant-MLX-2bit](https://huggingface.co/majentik/gemma-4-31b-it-turboquant-mlx-2bit) | mlx-lm | ~9.9 GB | Apple Silicon, smallest |
| [TurboQuant-MLX-4bit](https://huggingface.co/majentik/gemma-4-31b-it-turboquant-mlx-4bit) | mlx-lm | ~19 GB | Apple Silicon balanced |
| [TurboQuant-MLX-8bit](https://huggingface.co/majentik/gemma-4-31b-it-turboquant-mlx-8bit) | mlx-lm | ~37 GB | Apple Silicon reference |