majentik commited on
Commit
50287a7
·
verified ·
1 Parent(s): cf39760

docs: warn that the turboquant fork predates gemma4 support (reported in gemma-4-26B-A4B-RotorQuant#1)

Browse files
Files changed (1) hide show
  1. README.md +4 -0
README.md CHANGED
@@ -14,6 +14,10 @@ library_name: gguf
14
  pipeline_tag: image-text-to-text
15
  ---
16
 
 
 
 
 
17
  # gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M
18
 
19
  GGUF Q4_K_M weight-quantized variant of [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) optimised for use with **RotorQuant** KV cache compression via a dedicated llama.cpp fork.
 
14
  pipeline_tag: image-text-to-text
15
  ---
16
 
17
+ > [!WARNING]
18
+ > **Fork compatibility (2026-07-07):** the `llama-cpp-turboquant` fork is currently based on a llama.cpp revision that **predates `gemma4` architecture support** — it fails with `unknown model architecture: 'gemma4'` and cannot run this model at all. Until the fork rebases, use **mainline llama.cpp** (which loads this GGUF fine with standard KV-cache types); the RotorQuant/TurboQuant KV-cache options are not usable with gemma-4 yet.
19
+ <!-- gemma4-fork-note -->
20
+
21
  # gemma-4-31B-it-RotorQuant-GGUF-Q4_K_M
22
 
23
  GGUF Q4_K_M weight-quantized variant of [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) optimised for use with **RotorQuant** KV cache compression via a dedicated llama.cpp fork.