---
license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
library_name: transformers
base_model:
- LilaRest/gemma-4-31B-it-NVFP4-turbo
pipeline_tag: text-generation
tags:
- gemma4
- gemma-4-31b-it
- nvfp4
- modelopt
- vllm
- quantized
- nvidia
- lighthouse
model-index:
- name: gemma-4-31B-it-NVFP4-turbo
results:
- task:
type: text-generation
dataset:
name: GPQA Diamond
type: Idavidrein/gpqa
config: gpqa_diamond
metrics:
- name: Accuracy
type: accuracy
value: 72.73
- task:
type: text-generation
dataset:
name: MMLU Pro
type: TIGER-Lab/MMLU-Pro
metrics:
- name: Accuracy
type: accuracy
value: 83.93
---
⚡ Gemma 4 31B IT NVFP4 Turbo GGUF
Requires [ggml-org/llama.cpp#21971](https://github.com/ggml-org/llama.cpp/pull/21971)
A repackaged [nvidia/Gemma-4-31B-IT-NVFP4](https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4) that is **68% smaller** in GPU memory and **~2.5× faster** than the [base model](https://huggingface.co/google/gemma-4-31B-it), while retaining **nearly identical quality** (1-3% loss). Fits on a *single* RTX 5090 (🎉).
## Approach
Three changes were made:
1. **Quantized** all self-attention weights from BF16 → FP4 (RTN, group_size=16, matching modelopt NVFP4 format)
2. **Updated** architecture to `Gemma4ForCausalLM` and quantization config accordingly
3. **Stripped** the vision and audio encoder
Everything else is untouched — MLP layers keep NVIDIA's calibrated FP4, `embed_tokens` stays BF16, all norms preserved, so we retain all the [nvidia/Gemma-4-31B-IT-NVFP4](https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4) optimizations.
#### Why RTN didn't hurt quality
RTN (Round-To-Nearest) is the simplest quantization method — no calibration data, fully reproducible. It worked here because:
- FP4 with group_size=16 and per-group scaling preserves relative weight distributions well
- Self-attention weights tend to be normally distributed near zero, where the FP4 grid has finest resolution (0, 0.5, 1.0, 1.5)
- MLP layers (more sensitive to quantization) keep NVIDIA's calibrated FP4
- `embed_tokens` stays BF16, preventing noise from propagating through all layers
## License
Apache 2.0 — same as the [base model](https://ai.google.dev/gemma/docs/gemma_4_license).
## Credits
- [Google DeepMind](https://deepmind.google/models/gemma/) for Gemma 4
- [NVIDIA](https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4) for the modelopt NVFP4 checkpoint