--- base_model: - unsloth/Meta-Llama-3.1-8B license: llama3.1 pipeline_tag: text-generation tags: - prismquant - quantization - gptq - w4a4kv4 - custom-code --- # PrismQuant-Llama-3.1-8B [Paper](https://arxiv.org/abs/2609.32429) · [Code](https://github.com/ForeverBlue816/PrismQuant) · [Loading guide](https://github.com/ForeverBlue816/PrismQuant/blob/main/docs/models.md) Official checkpoints for **PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers**. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups. ## Models | Model key | Base model | Quantization | | --- | --- | --- | | `llama31_8b` | [unsloth/Meta-Llama-3.1-8B](https://huggingface.co/unsloth/Meta-Llama-3.1-8B) | W4A4KV4 | ## Quick start ```bash git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git cd PrismQuant pip install -e . ``` ```python from prismquant import load_model loaded = load_model("llama31_8b") print(loaded.generate("The key idea behind quantization is", max_new_tokens=64)) ``` The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion. ## Checkpoint format The GPTQ INT4 values are stored as **dequantized floating-point tensors**. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model. Reference compute: **bfloat16**. Budget for the base-model state, activations and KV cache as well as the downloaded tensors. `config.json` declares the PrismQuant artifact format and default models. `checkpoints/` contains decoder weights; `rotations/` contains the factors required by the loader. ## License **Built with Llama.** The upstream Llama Community License applies. See [LICENSE](LICENSE), [NOTICE](NOTICE) and [Acceptable Use Policy](USE_POLICY.md).