ForeverBlue commited on
Commit
fc812e6
·
verified ·
1 Parent(s): d5acd2d

Publish PrismQuant paper, concise model card, and loader manifest

Browse files
Files changed (2) hide show
  1. README.md +15 -34
  2. config.json +13 -0
README.md CHANGED
@@ -9,18 +9,20 @@ tags:
9
  - gptq
10
  - w4a4kv4
11
  - custom-code
12
- - research
13
  ---
14
 
15
- # Llama 3.1 — PrismQuant
16
 
17
- Released checkpoints for [PrismQuant](https://github.com/ForeverBlue816/PrismQuant): quantizer-aware rotations for low-bit language models.
18
 
19
- **Base models:** [unsloth/Meta-Llama-3.1-8B](https://huggingface.co/unsloth/Meta-Llama-3.1-8B).
20
 
21
- **arXiv:** coming soon.
22
 
23
- **Built with Llama.** These derivatives retain the Llama 3.1 Community License. See [LICENSE](LICENSE), [NOTICE](NOTICE) and [Acceptable Use Policy](USE_POLICY.md).
 
 
24
 
25
  ## Quick start
26
 
@@ -28,7 +30,6 @@ Released checkpoints for [PrismQuant](https://github.com/ForeverBlue816/PrismQua
28
  git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
29
  cd PrismQuant
30
  pip install -e .
31
- prismquant models
32
  ```
33
 
34
  ```python
@@ -38,36 +39,16 @@ loaded = load_model("llama31_8b")
38
  print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))
39
  ```
40
 
41
- The loader selectively downloads one checkpoint and its required rotation factors, pins its revision, and obtains the matching base model/tokenizer. Defaults use PrismQuant `k=max`, seed 0. Use `prismquant models` to select another exact checkpoint name. See the [loading guide](https://github.com/ForeverBlue816/PrismQuant/blob/main/docs/models.md).
42
 
43
- ## Format and limitations
44
 
45
- These are **dequantized GPTQ INT4 weights stored in floating-point tensors**, not packed INT4 weights or standalone Transformers checkpoints. Do not pass this repository directly to `AutoModelForCausalLM.from_pretrained`. The PrismQuant loader restores unquantized state from the base model, loads the saved decoder linears, and installs the online rotations and activation/KV quantizers.
46
 
47
- The reference runtime retains floating-point weight and KV storage. It uses the original bf16 containers for Llama 3B/8B and fp32 for Qwen3/70B. W4A4KV4 names the simulated numerical quantization, not this artifact's disk size or serving memory. Dedicated packed-kernel deployment benchmarks are documented separately in the code repository.
48
 
49
- These are base-model completions rather than chat-tuned assistants. The original evaluation protocols, seeds, caveats and negative findings remain in the [research archive](https://github.com/ForeverBlue816/PrismQuant/tree/research-archive-2026-09-25). Quantization does not remove upstream model limitations.
50
 
51
- ## Contents
52
 
53
- - `checkpoints/<model>/<checkpoint>/layer_XX.pt`: the folded, GPTQ-quantized decoder linear tensors.
54
- - `DONE.json` and `gptq_audit.csv`: quantization settings and provenance.
55
- - `rotations/<model>/...`: the calibrated factors used by the runtime.
56
-
57
- The `nar` strings in filenames are the original internal name of PrismQuant and are retained for compatibility. Existing weight files are unchanged by this public-release update. The two smaller Llama repositories additionally include the previously omitted `k=8` down-projection factors.
58
-
59
- ## Available checkpoints
60
-
61
- Download sizes below cover one variant and its required factors, excluding the separately downloaded base model.
62
-
63
- | Model key | Checkpoint | Weight protocol | Download (GB) |
64
- | --- | --- | --- | --- |
65
- | `llama31_8b` | `gptq_hadamard_seed0` | default | 13.96 |
66
- | `llama31_8b` | `gptq_hadamard_seed1` | default | 13.96 |
67
- | `llama31_8b` | `gptq_hadamard_seed2` | default | 13.96 |
68
- | `llama31_8b` | `gptq_nar_k8_seed0` | default | 13.98 |
69
- | `llama31_8b` | `gptq_nar_k8_seed1` | default | 13.98 |
70
- | `llama31_8b` | `gptq_nar_k8_seed2` | default | 13.98 |
71
- | `llama31_8b` | `gptq_nar_kmax_seed0` | default | 14.17 |
72
- | `llama31_8b` | `gptq_nar_kmax_seed1` | default | 14.17 |
73
- | `llama31_8b` | `gptq_nar_kmax_seed2` | default | 14.17 |
 
9
  - gptq
10
  - w4a4kv4
11
  - custom-code
12
+ - arxiv:2609.32429
13
  ---
14
 
15
+ # PrismQuant-Llama-3.1-8B
16
 
17
+ [Paper](https://arxiv.org/abs/2609.32429) · [Code](https://github.com/ForeverBlue816/PrismQuant) · [Loading guide](https://github.com/ForeverBlue816/PrismQuant/blob/main/docs/models.md)
18
 
19
+ Official checkpoints for **PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers**. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.
20
 
21
+ ## Models
22
 
23
+ | Model key | Base model | Quantization |
24
+ | --- | --- | --- |
25
+ | `llama31_8b` | [unsloth/Meta-Llama-3.1-8B](https://huggingface.co/unsloth/Meta-Llama-3.1-8B) | W4A4KV4 |
26
 
27
  ## Quick start
28
 
 
30
  git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
31
  cd PrismQuant
32
  pip install -e .
 
33
  ```
34
 
35
  ```python
 
39
  print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))
40
  ```
41
 
42
+ The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.
43
 
44
+ ## Checkpoint format
45
 
46
+ The GPTQ INT4 values are stored as **dequantized floating-point tensors**. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.
47
 
48
+ Reference compute: **bfloat16**. Budget for the base-model state, activations and KV cache as well as the downloaded tensors.
49
 
50
+ `config.json` declares the PrismQuant artifact format and default models. `checkpoints/` contains decoder weights; `rotations/` contains the factors required by the loader.
51
 
52
+ ## License
53
 
54
+ **Built with Llama.** The upstream Llama Community License applies. See [LICENSE](LICENSE), [NOTICE](NOTICE) and [Acceptable Use Policy](USE_POLICY.md).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "prismquant-reference-v1",
3
+ "paper": "https://arxiv.org/abs/2609.32429",
4
+ "models": {
5
+ "llama31_8b": {
6
+ "base_model": "unsloth/Meta-Llama-3.1-8B",
7
+ "default_checkpoint": "gptq_nar_kmax_seed0",
8
+ "architecture": "llama",
9
+ "quantization": "W4A4KV4",
10
+ "storage_dtype": "bfloat16"
11
+ }
12
+ }
13
+ }