ForeverBlue commited on
Commit
48352eb
·
verified ·
1 Parent(s): 5195fee

Publish PrismQuant paper, concise model card, and loader manifest

Browse files
Files changed (2) hide show
  1. README.md +18 -37
  2. config.json +34 -0
README.md CHANGED
@@ -12,18 +12,23 @@ tags:
12
  - gptq
13
  - w4a4kv4
14
  - custom-code
15
- - research
16
  ---
17
 
18
- # Qwen3 Base — PrismQuant
19
 
20
- Released checkpoints for [PrismQuant](https://github.com/ForeverBlue816/PrismQuant): quantizer-aware rotations for low-bit language models.
21
 
22
- **Base models:** [Qwen/Qwen3-0.6B-Base](https://huggingface.co/Qwen/Qwen3-0.6B-Base), [Qwen/Qwen3-1.7B-Base](https://huggingface.co/Qwen/Qwen3-1.7B-Base), [Qwen/Qwen3-4B-Base](https://huggingface.co/Qwen/Qwen3-4B-Base), [Qwen/Qwen3-8B-Base](https://huggingface.co/Qwen/Qwen3-8B-Base).
23
 
24
- **arXiv:** coming soon.
25
 
26
- Model weights retain [Apache-2.0](LICENSE).
 
 
 
 
 
27
 
28
  ## Quick start
29
 
@@ -31,7 +36,6 @@ Model weights retain [Apache-2.0](LICENSE).
31
  git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
32
  cd PrismQuant
33
  pip install -e .
34
- prismquant models
35
  ```
36
 
37
  ```python
@@ -41,39 +45,16 @@ loaded = load_model("qwen3_0.6b_base")
41
  print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))
42
  ```
43
 
44
- The loader selectively downloads one checkpoint and its required rotation factors, pins its revision, and obtains the matching base model/tokenizer. Defaults use PrismQuant `k=max`, seed 0. Use `prismquant models` to select another exact checkpoint name. See the [loading guide](https://github.com/ForeverBlue816/PrismQuant/blob/main/docs/models.md).
45
 
46
- ## Format and limitations
47
 
48
- These are **dequantized GPTQ INT4 weights stored in floating-point tensors**, not packed INT4 weights or standalone Transformers checkpoints. Do not pass this repository directly to `AutoModelForCausalLM.from_pretrained`. The PrismQuant loader restores unquantized state from the base model, loads the saved decoder linears, and installs the online rotations and activation/KV quantizers.
49
 
50
- The reference runtime retains floating-point weight and KV storage. It uses the original bf16 containers for Llama 3B/8B and fp32 for Qwen3/70B. W4A4KV4 names the simulated numerical quantization, not this artifact's disk size or serving memory. Dedicated packed-kernel deployment benchmarks are documented separately in the code repository.
51
 
52
- These are base-model completions rather than chat-tuned assistants. The original evaluation protocols, seeds, caveats and negative findings remain in the [research archive](https://github.com/ForeverBlue816/PrismQuant/tree/research-archive-2026-09-25). Quantization does not remove upstream model limitations.
53
 
54
- ## Contents
55
 
56
- - `checkpoints/<model>/<checkpoint>/layer_XX.pt`: the folded, GPTQ-quantized decoder linear tensors.
57
- - `DONE.json` and `gptq_audit.csv`: quantization settings and provenance.
58
- - `rotations/<model>/...`: the calibrated factors used by the runtime.
59
-
60
- The `nar` strings in filenames are the original internal name of PrismQuant and are retained for compatibility. Existing weight files are unchanged by this public-release update. The two smaller Llama repositories additionally include the previously omitted `k=8` down-projection factors.
61
-
62
- ## Available checkpoints
63
-
64
- Download sizes below cover one variant and its required factors, excluding the separately downloaded base model.
65
-
66
- | Model key | Checkpoint | Weight protocol | Download (GB) |
67
- | --- | --- | --- | --- |
68
- | `qwen3_0.6b_base` | `gptq_hadamard_seed0_g128_asym` | g128_asym | 1.76 |
69
- | `qwen3_0.6b_base` | `gptq_nar_k8_seed0_g128_asym` | g128_asym | 1.77 |
70
- | `qwen3_0.6b_base` | `gptq_nar_kmax_seed0_g128_asym` | g128_asym | 1.77 |
71
- | `qwen3_1.7b_base` | `gptq_hadamard_seed0_g128_asym` | g128_asym | 5.64 |
72
- | `qwen3_1.7b_base` | `gptq_nar_k8_seed0_g128_asym` | g128_asym | 5.65 |
73
- | `qwen3_1.7b_base` | `gptq_nar_kmax_seed0_g128_asym` | g128_asym | 5.67 |
74
- | `qwen3_4b_base` | `gptq_hadamard_seed0_g128_asym` | g128_asym | 14.53 |
75
- | `qwen3_4b_base` | `gptq_nar_k8_seed0_g128_asym` | g128_asym | 14.55 |
76
- | `qwen3_4b_base` | `gptq_nar_kmax_seed0_g128_asym` | g128_asym | 14.65 |
77
- | `qwen3_8b_base` | `gptq_hadamard_seed0_g128_asym` | g128_asym | 27.78 |
78
- | `qwen3_8b_base` | `gptq_nar_k8_seed0_g128_asym` | g128_asym | 27.80 |
79
- | `qwen3_8b_base` | `gptq_nar_kmax_seed0_g128_asym` | g128_asym | 27.96 |
 
12
  - gptq
13
  - w4a4kv4
14
  - custom-code
15
+ - arxiv:2609.32429
16
  ---
17
 
18
+ # PrismQuant-Qwen3-Base
19
 
20
+ [Paper](https://arxiv.org/abs/2609.32429) · [Code](https://github.com/ForeverBlue816/PrismQuant) · [Loading guide](https://github.com/ForeverBlue816/PrismQuant/blob/main/docs/models.md)
21
 
22
+ Official checkpoints for **PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers**. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.
23
 
24
+ ## Models
25
 
26
+ | Model key | Base model | Quantization |
27
+ | --- | --- | --- |
28
+ | `qwen3_0.6b_base` | [Qwen/Qwen3-0.6B-Base](https://huggingface.co/Qwen/Qwen3-0.6B-Base) | W4A4KV4 |
29
+ | `qwen3_1.7b_base` | [Qwen/Qwen3-1.7B-Base](https://huggingface.co/Qwen/Qwen3-1.7B-Base) | W4A4KV4 |
30
+ | `qwen3_4b_base` | [Qwen/Qwen3-4B-Base](https://huggingface.co/Qwen/Qwen3-4B-Base) | W4A4KV4 |
31
+ | `qwen3_8b_base` | [Qwen/Qwen3-8B-Base](https://huggingface.co/Qwen/Qwen3-8B-Base) | W4A4KV4 |
32
 
33
  ## Quick start
34
 
 
36
  git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
37
  cd PrismQuant
38
  pip install -e .
 
39
  ```
40
 
41
  ```python
 
45
  print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))
46
  ```
47
 
48
+ The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.
49
 
50
+ ## Checkpoint format
51
 
52
+ The GPTQ INT4 values are stored as **dequantized floating-point tensors**. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.
53
 
54
+ Reference compute: **float32**. Budget for the base-model state, activations and KV cache as well as the downloaded tensors.
55
 
56
+ `config.json` declares the PrismQuant artifact format and default models. `checkpoints/` contains decoder weights; `rotations/` contains the factors required by the loader.
57
 
58
+ ## License
59
 
60
+ Qwen-derived weights retain [Apache-2.0](LICENSE).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
config.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "prismquant-reference-v1",
3
+ "paper": "https://arxiv.org/abs/2609.32429",
4
+ "models": {
5
+ "qwen3_0.6b_base": {
6
+ "base_model": "Qwen/Qwen3-0.6B-Base",
7
+ "default_checkpoint": "gptq_nar_kmax_seed0_g128_asym",
8
+ "architecture": "qwen3",
9
+ "quantization": "W4A4KV4",
10
+ "storage_dtype": "float32"
11
+ },
12
+ "qwen3_1.7b_base": {
13
+ "base_model": "Qwen/Qwen3-1.7B-Base",
14
+ "default_checkpoint": "gptq_nar_kmax_seed0_g128_asym",
15
+ "architecture": "qwen3",
16
+ "quantization": "W4A4KV4",
17
+ "storage_dtype": "float32"
18
+ },
19
+ "qwen3_4b_base": {
20
+ "base_model": "Qwen/Qwen3-4B-Base",
21
+ "default_checkpoint": "gptq_nar_kmax_seed0_g128_asym",
22
+ "architecture": "qwen3",
23
+ "quantization": "W4A4KV4",
24
+ "storage_dtype": "float32"
25
+ },
26
+ "qwen3_8b_base": {
27
+ "base_model": "Qwen/Qwen3-8B-Base",
28
+ "default_checkpoint": "gptq_nar_kmax_seed0_g128_asym",
29
+ "architecture": "qwen3",
30
+ "quantization": "W4A4KV4",
31
+ "storage_dtype": "float32"
32
+ }
33
+ }
34
+ }