File size: 2,044 Bytes
51c82f0 d5acd2d d94ada1 d5acd2d 51c82f0 fc812e6 51c82f0 fc812e6 51c82f0 fc812e6 51c82f0 fc812e6 51c82f0 fc812e6 51c82f0 d5acd2d 51c82f0 d5acd2d 51c82f0 d5acd2d 51c82f0 d5acd2d 51c82f0 fc812e6 d5acd2d fc812e6 d5acd2d fc812e6 d5acd2d fc812e6 d5acd2d fc812e6 d5acd2d fc812e6 d5acd2d d94ada1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 | ---
base_model:
- unsloth/Meta-Llama-3.1-8B
license: llama3.1
pipeline_tag: text-generation
tags:
- prismquant
- quantization
- gptq
- w4a4kv4
- custom-code
---
# PrismQuant-Llama-3.1-8B
[Paper](https://arxiv.org/abs/2609.32429) · [Code](https://github.com/ForeverBlue816/PrismQuant) · [Loading guide](https://github.com/ForeverBlue816/PrismQuant/blob/main/docs/models.md)
Official checkpoints for **PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers**. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.
## Models
| Model key | Base model | Quantization |
| --- | --- | --- |
| `llama31_8b` | [unsloth/Meta-Llama-3.1-8B](https://huggingface.co/unsloth/Meta-Llama-3.1-8B) | W4A4KV4 |
## Quick start
```bash
git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
cd PrismQuant
pip install -e .
```
```python
from prismquant import load_model
loaded = load_model("llama31_8b")
print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))
```
The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.
## Checkpoint format
The GPTQ INT4 values are stored as **dequantized floating-point tensors**. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.
Reference compute: **bfloat16**. Budget for the base-model state, activations and KV cache as well as the downloaded tensors.
`config.json` declares the PrismQuant artifact format and default models. `checkpoints/` contains decoder weights; `rotations/` contains the factors required by the loader.
## License
**Built with Llama.** The upstream Llama Community License applies. See [LICENSE](LICENSE), [NOTICE](NOTICE) and [Acceptable Use Policy](USE_POLICY.md). |