gemma-4-E2B-it-uncensored — Q4_0 GGUF

Q4_0 quantization of TrevorJS/gemma-4-E2B-it-uncensored, made for fast CPU inference on ARM phones (llama.cpp repacks Q4_0 weights for the NEON/DOTPROD kernels).

This is a modified version of the original work: the weights were re-quantized. Nothing else was changed.

  • Source: gemma-4-E2B-it-uncensored-Q8_0.gguf from TrevorJS/gemma-4-E2B-it-uncensored-GGUF
  • Tool: llama-quantize --allow-requantize … Q4_0 (llama.cpp b11228)
  • Result: transformer layers in q4_0, token_embd and per_layer_token_embd in q6_K
  • File: gemma-4-E2B-it-uncensored-Q4_0.gguf, 3,360,154,144 bytes
  • SHA-256: 06a0d541e0aba58c8bfb5ea885493b75c291f176fb9ea154099cad7b80f87446

Measured on a Kirin 980 (4 threads, CPU only, llama.rn 0.12.9): prompt processing ~50 tok/s vs ~37 tok/s for the author's Q4_K_M; generation ~10.6 tok/s vs ~9.5 tok/s (JSON-constrained).

License

Apache License 2.0, same as the original model. All credit for the model and the uncensoring method goes to TrevorJS and, for the base model, Google (Gemma 4).

Downloads last month
234
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for n1c0c4b/gemma-4-E2B-it-uncensored-Q4_0-GGUF

Quantized
(17)
this model