--- base_model: llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic base_model_relation: quantized quantized_by: DBMe library_name: exllamav3 pipeline_tag: text-generation license: apache-2.0 tags: - exl3 - exllamav3 - quantized - text-generation --- # gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1 EXL3 (ExLlamaV3) quantizations of [llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic](https://huggingface.co/llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic). All credit for the original model goes to the original authors. ## ๐Ÿ“Š Available Quantizations & VRAM The model weights are stored in separate branches. **Please switch to a branch to download.** *Note: VRAM estimates include PyTorch context overhead (~0.8GB) and assume an unquantized FP16 KV cache.* | Target BPW | Head BPW | Branch (Download Link) | WikiText-2 PPL (512 ctx)ยน | 2K ctx | 4K ctx | 8K ctx | 16K ctx | 32K ctx | |---|---|---|---|---|---|---|---|---| | 3.5 | h6 | [3.5bpw_h6](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/tree/3.5bpw_h6) | 2335.9614 | N/A | N/A | N/A | N/A | N/A | | 3.75 | h6 | [3.75bpw_h6](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/tree/3.75bpw_h6) | 2312.4667 | N/A | N/A | N/A | N/A | N/A | | 4.0 | h6 | [4.0bpw_h6](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/tree/4.0bpw_h6) | 2097.6864 | N/A | N/A | N/A | N/A | N/A | | 5.0 | h6 | [5.0bpw_h6](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/tree/5.0bpw_h6) | 1995.7812 | N/A | N/A | N/A | N/A | N/A | | 6.0 | h6 | [6.0bpw_h6](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/tree/6.0bpw_h6) | 1995.3135 | N/A | N/A | N/A | N/A | N/A | | 8.0 | h8 | [8.0bpw_h8](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/tree/8.0bpw_h8) | 1985.1562 | N/A | N/A | N/A | N/A | N/A | ยน *Evaluated against WikiText-2 with ExLlamaV3 using a strided 512-token context window (-c 512) in llama.cpp parity mode (-g). Lower is better.* *(Higher BPW = higher quality, lower BPW = fits in less VRAM).* ## ๐Ÿ“ฅ How to Download It's recommended to use the `huggingface-cli` to download specific branches. *(Do not use `git clone` as it will download all branches!)* Ensure you have the CLI installed: ```bash pip install -U "huggingface_hub[cli]" ``` Download a specific branch (e.g., `5.0bpw_h6`): ```bash # Example: Downloading the 5.0bpw_h6 branch huggingface-cli download DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1 --revision 5.0bpw_h6 --local-dir gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1-5.0bpw_h6 ``` ## ๐Ÿ’ป Supported Engines These models are highly optimized for modern GPUs and can be run using: * **[TabbyAPI](https://github.com/theroyallab/tabbyAPI):** A fast, OpenAI-compatible API server. *(Set `model_name` to the local folder name you downloaded the branch into)* * **[Text-Generation-WebUI](https://github.com/oobabooga/text-generation-webui):** A local web interface. *(Select the `exllamav3` loader)* * **[ExLlamaV3 (Native)](https://github.com/turboderp-org/exllamav3):** Python library for custom integration. ### ๐Ÿ“ˆ Perplexity Degradation Curve *(Lower is better)* ![Perplexity Graph](https://huggingface.co/DBMe/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-exl3-mul1/resolve/main/metrics_graph.png?v=ee7a7541)
โš™๏ธ Advanced: Quantization Environment & Settings ### ๐Ÿ”ฌ Quantization Settings - **Codebook:** mul1 - **Output Scales:** always - **Calibration Rows:** 250 - **Calibration Cols:** 2048 - **Calibration Dataset:** ExLlamaV3 Default (Wiki/C4/Code) - **High Quality (HQ) Mode:** False - **ExLlamaV3:** `1.0.0` (Commit: `cb7f2e9`) - **Hardware:** `NVIDIA RTX PRO 6000 Blackwell Server Edition`