PrismQuant-Llama-3.1-8B

Paper · Code · Loading guide

Official checkpoints for PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.

Models

Model key Base model Quantization
llama31_8b unsloth/Meta-Llama-3.1-8B W4A4KV4

Quick start

git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
cd PrismQuant
pip install -e .
from prismquant import load_model

loaded = load_model("llama31_8b")
print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))

The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.

Checkpoint format

The GPTQ INT4 values are stored as dequantized floating-point tensors. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.

Reference compute: bfloat16. Budget for the base-model state, activations and KV cache as well as the downloaded tensors.

config.json declares the PrismQuant artifact format and default models. checkpoints/ contains decoder weights; rotations/ contains the factors required by the loader.

License

Built with Llama. The upstream Llama Community License applies. See LICENSE, NOTICE and Acceptable Use Policy.

Downloads last month
353
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ForeverBlue/PrismQuant-Llama-3.1-8B

Finetuned
(300)
this model

Collection including ForeverBlue/PrismQuant-Llama-3.1-8B

Paper for ForeverBlue/PrismQuant-Llama-3.1-8B