majentik commited on
Commit
90f7b21
·
verified ·
1 Parent(s): bec001f

docs: Tier 2 polish — variant matrix + quant trade-off

Browse files
Files changed (1) hide show
  1. README.md +40 -0
README.md CHANGED
@@ -157,3 +157,43 @@ For VRAM-constrained setups, standard q8_0 KV cache quantization already halves
157
  - [TurboQuant paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
158
  - [Base model: google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it)
159
  - [gemma-4-31B-it announcement](https://blog.google/technology/developers/gemma-4/)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
157
  - [TurboQuant paper (arXiv: 2504.19874)](https://arxiv.org/abs/2504.19874)
158
  - [Base model: google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it)
159
  - [gemma-4-31B-it announcement](https://blog.google/technology/developers/gemma-4/)
160
+
161
+ ## Quant trade-off (GGUF lane)
162
+
163
+ | Quant | Approx size | Use case | Recommendation |
164
+ |---|---|---|---|
165
+ | Q2_K | ~17 GB | Lossy, low-RAM CPU/edge | Resource-constrained inference |
166
+ | Q3_K_M | ~19 GB | Smaller-than-Q4, modest quality drop | Edge devices with ~16 GB RAM |
167
+ | IQ4_XS | ~16 GB | Importance-quant 4-bit, smaller than Q4_K_M | Best size/quality at 4-bit |
168
+ | **Q4_K_M** | ~23 GB | Balanced default | **Recommended for most users** |
169
+ | Q5_K_M | ~24 GB | Higher fidelity than Q4 | Quality-sensitive applications |
170
+ | Q6_K | ~28 GB | Approaching FP16 quality | High-fidelity CPU/edge |
171
+ | Q8_0 | ~32 GB | Near-lossless reference | Fidelity-critical work |
172
+ | MXFP4_MOE | ~17 GB | Microscaling FP4 (MoE-aware) | vLLM / transformers users |
173
+
174
+ (Current variant — **Q4_K_M** — is bolded.)
175
+
176
+ ## Variants in this family
177
+
178
+ (Showing 18 sibling variants under `majentik/gemma4-31b-it-*`. The current variant — `RotorQuant-GGUF-Q4_K_M` — is **bolded**.)
179
+
180
+ | Variant | Runtime | Approx size | Use case |
181
+ |---|---|---|---|
182
+ | [RotorQuant](https://huggingface.co/majentik/gemma4-31b-it-rotorquant) | runtime modifier | n/a | KV-cache root (weight-agnostic) |
183
+ | [RotorQuant-AWQ-4bit](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-awq-4bit) | transformers | ~19 GB | GPU 4-bit (AutoAWQ) |
184
+ | [RotorQuant-AWQ-8bit](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-awq-8bit) | transformers | ~34 GB | GPU 8-bit (AutoAWQ) |
185
+ | [RotorQuant-GGUF-IQ4_XS](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-gguf-IQ4_XS) | llama.cpp | ~27 GB | Lossy 4-bit, low-RAM CPU/edge |
186
+ | [RotorQuant-GGUF-Q2_K](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-gguf-Q2_K) | llama.cpp | ~19 GB | Lossy, low-RAM CPU/edge |
187
+ | [RotorQuant-GGUF-Q3_K_M](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-gguf-Q3_K_M) | llama.cpp | ~24 GB | Smaller 3-bit, CPU-friendly |
188
+ | **RotorQuant-GGUF-Q4_K_M** | llama.cpp | ~34 GB | Balanced default |
189
+ | [RotorQuant-GGUF-Q5_K_M](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-gguf-Q5_K_M) | llama.cpp | ~41 GB | Higher fidelity, more RAM |
190
+ | [RotorQuant-GGUF-Q8_0](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-gguf-Q8_0) | llama.cpp | ~65 GB | Near-lossless reference |
191
+ | [RotorQuant-MLX-2bit](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-mlx-2bit) | mlx-lm | ~9.9 GB | Apple Silicon, smallest |
192
+ | [RotorQuant-MLX-4bit](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-mlx-4bit) | mlx-lm | ~19 GB | Apple Silicon balanced |
193
+ | [RotorQuant-MLX-8bit](https://huggingface.co/majentik/gemma4-31b-it-rotorquant-mlx-8bit) | mlx-lm | ~37 GB | Apple Silicon reference |
194
+ | [TurboQuant](https://huggingface.co/majentik/gemma4-31b-it-turboquant) | runtime modifier | n/a | KV-cache root (weight-agnostic) |
195
+ | [TurboQuant-AWQ-4bit](https://huggingface.co/majentik/gemma4-31b-it-turboquant-awq-4bit) | transformers | ~19 GB | GPU 4-bit (AutoAWQ) |
196
+ | [TurboQuant-AWQ-8bit](https://huggingface.co/majentik/gemma4-31b-it-turboquant-awq-8bit) | transformers | ~34 GB | GPU 8-bit (AutoAWQ) |
197
+ | [TurboQuant-MLX-2bit](https://huggingface.co/majentik/gemma4-31b-it-turboquant-mlx-2bit) | mlx-lm | ~9.9 GB | Apple Silicon, smallest |
198
+ | [TurboQuant-MLX-4bit](https://huggingface.co/majentik/gemma4-31b-it-turboquant-mlx-4bit) | mlx-lm | ~19 GB | Apple Silicon balanced |
199
+ | [TurboQuant-MLX-8bit](https://huggingface.co/majentik/gemma4-31b-it-turboquant-mlx-8bit) | mlx-lm | ~37 GB | Apple Silicon reference |