Llama-2-70b-hf — KronQ W2A16 (packed int2)

Paper: arXiv:2607.07964 · Code: GitHub

Llama-2-70b-hf quantized to 2-bit weights / 16-bit activations with KronQ (Kronecker-factored Hessian quantization). Weights are stored packed int2 (~17 GB); a fused dequant + bidirectional-incoherence (BiIP) CUDA kernel unpacks them on the fly.

Results (WikiText-2, seqlen 2048)

Perplexity: 5.14

Zero-shot accuracy:

PIQA ARC-E ARC-C HellaSwag WinoGrande BoolQ OBQA Average
77.37 75.93 46.93 71.98 74.98 79.69 43.20 67.15

(lm-evaluation-harness, 0-shot. acc_norm for PIQA/HellaSwag/ARC/OBQA, acc for WinoGrande/BoolQ.)

Usage

KronQ-packed checkpoint. Load with the KronQ runtime:

python eval_pretrained.py meta-llama/Llama-2-70b-hf donghyunli/Llama-2-70b-KronQ-W2A16 --ppl --zs

Recipe

Per-channel asymmetric W2, weight-only (a_bits=16), --alpha 0.25, BiIP, act_order, raw H_G. 128 WikiText-2 calibration sequences.

License

Derivative of Llama-2-70b-hf — llama2 license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
18B params
Tensor type
F32
·
F16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for donghyunli/Llama-2-70b-KronQ-W2A16

Finetuned
(36)
this model

Paper for donghyunli/Llama-2-70b-KronQ-W2A16