--- license: apache-2.0 license_link: https://ai.google.dev/gemma/docs/gemma_4_license library_name: transformers base_model: - LilaRest/gemma-4-31B-it-NVFP4-turbo pipeline_tag: text-generation tags: - gemma4 - gemma-4-31b-it - nvfp4 - modelopt - vllm - quantized - nvidia - lighthouse model-index: - name: gemma-4-31B-it-NVFP4-turbo results: - task: type: text-generation dataset: name: GPQA Diamond type: Idavidrein/gpqa config: gpqa_diamond metrics: - name: Accuracy type: accuracy value: 72.73 - task: type: text-generation dataset: name: MMLU Pro type: TIGER-Lab/MMLU-Pro metrics: - name: Accuracy type: accuracy value: 83.93 ---

⚡ Gemma 4 31B IT NVFP4 Turbo GGUF

Requires [ggml-org/llama.cpp#21971](https://github.com/ggml-org/llama.cpp/pull/21971) A repackaged [nvidia/Gemma-4-31B-IT-NVFP4](https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4) that is **68% smaller** in GPU memory and **~2.5× faster** than the [base model](https://huggingface.co/google/gemma-4-31B-it), while retaining **nearly identical quality** (1-3% loss). Fits on a *single* RTX 5090 (🎉). ## Approach Three changes were made: 1. **Quantized** all self-attention weights from BF16 → FP4 (RTN, group_size=16, matching modelopt NVFP4 format) 2. **Updated** architecture to `Gemma4ForCausalLM` and quantization config accordingly 3. **Stripped** the vision and audio encoder Everything else is untouched — MLP layers keep NVIDIA's calibrated FP4, `embed_tokens` stays BF16, all norms preserved, so we retain all the [nvidia/Gemma-4-31B-IT-NVFP4](https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4) optimizations. #### Why RTN didn't hurt quality RTN (Round-To-Nearest) is the simplest quantization method — no calibration data, fully reproducible. It worked here because: - FP4 with group_size=16 and per-group scaling preserves relative weight distributions well - Self-attention weights tend to be normally distributed near zero, where the FP4 grid has finest resolution (0, 0.5, 1.0, 1.5) - MLP layers (more sensitive to quantization) keep NVIDIA's calibrated FP4 - `embed_tokens` stays BF16, preventing noise from propagating through all layers ## License Apache 2.0 — same as the [base model](https://ai.google.dev/gemma/docs/gemma_4_license). ## Credits - [Google DeepMind](https://deepmind.google/models/gemma/) for Gemma 4 - [NVIDIA](https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4) for the modelopt NVFP4 checkpoint