PrismQuant-Llama-3.1-70B

Paper · Code · Loading guide

Official checkpoints for PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.

Models

Model key Base model Quantization
llama31_70b unsloth/Meta-Llama-3.1-70B W4A4KV4

Quick start

git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
cd PrismQuant
pip install -e .
from prismquant import load_model

loaded = load_model("llama31_70b", device_map="balanced")
print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))

The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.

Checkpoint format

The GPTQ INT4 values are stored as dequantized floating-point tensors. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.

This fp32 reference requires roughly 280 GB for model parameters, plus runtime overhead. Use sufficient aggregate GPU memory and, if needed, max_memory={...} with balanced placement.

config.json declares the PrismQuant artifact format and default models. checkpoints/ contains decoder weights; rotations/ contains the factors required by the loader.

License

Built with Llama. The upstream Llama Community License applies. See LICENSE, NOTICE and Acceptable Use Policy.

Downloads last month
361
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ForeverBlue/PrismQuant-Llama-3.1-70B

Finetuned
(1)
this model

Collection including ForeverBlue/PrismQuant-Llama-3.1-70B

Paper for ForeverBlue/PrismQuant-Llama-3.1-70B