File size: 3,932 Bytes
702702a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5b0af84
702702a
 
 
 
 
b7bbd3b
702702a
b7bbd3b
 
702702a
6e7c07b
b7bbd3b
 
702702a
 
b7bbd3b
 
702702a
 
 
 
b7bbd3b
 
 
 
 
 
 
702702a
b7bbd3b
702702a
b7bbd3b
702702a
b7bbd3b
 
 
 
afd1084
 
 
 
 
b7bbd3b
 
 
c2a0369
b7bbd3b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
702702a
 
 
 
 
b7bbd3b
 
 
 
 
 
 
702702a
 
 
 
 
b7bbd3b
 
 
702702a
 
 
 
 
 
 
 
 
 
 
b7bbd3b
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
---
base_model: Qwen/Qwen3-1.7B
language:
- en
license: apache-2.0
pipeline_tag: text-generation
library_name: transformers
arxiv: 2509.22944
tags:
- quantized
- sinq
- efficient-inference
- qwen
- llm
- compression
base_model_relation: quantized
---

<p align="center">
  <img src="SINQ_GGUF_HF.png" alt="Logo" style="max-width: 80%; height: auto;">
</p>

<p align="center">πŸ™ <a href="https://github.com/huawei-csl/SINQ">Github</a>&nbsp;&nbsp; | &nbsp;&nbsp;πŸ“„ <a href="http://arxiv.org/abs/2509.22944">Paper</a></p>


# PreSINQ GGUF Quantized Qwen3-1.7B Model

This repository contains the official PreSINQ **GGUF-quantized** versions of the [`Qwen3-1.7B`](https://huggingface.co/Qwen/Qwen3-1.7B) model. For a detailed explanation of PreSINQ strategy please refer to the the official [SINQ](https://github.com/huawei-csl/SINQ) repository.
SINQ is a fast and high-quality quantization technique designed to significantly reduce Large Language Model size while preserving accuracy.

If you find this project useful, **please consider giving a ⭐ to the official [SINQ](https://github.com/huawei-csl/SINQ) repository**.

---

## Model Details

- **Model Name:** `Qwen3-1.7B-PreSINQ-GGUF`
- **Base Model:** [`Qwen/Qwen3-1.7B`](https://huggingface.co/Qwen/Qwen3-1.7B)
- **Task:** Text Generation
- **Framework:** PyTorch / Transformers
- **License:** [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0)
- **Quantized By:** *Huawei – Computing Systems Lab*

---

# How to Obtain the PreSINQ Model

The PreSINQ Qwen3-1.7B models are produced using the **PreSINQ GGUF script** available in the official [SINQ](https://github.com/huawei-csl/SINQ) repository.

The models provided here correspond to the best-performing configurations for each quantization type.

## πŸ“Š Best PreSINQ Quantization Results (Qwen3-1.7B)

Results below are measured on the **WikiText-2 test set**.

| Method | Bits | Size (GB) | Perplexity ↓ |
|----------|--------|------------|----------------|
| Baseline (FP16) | FP16 | 3.79 | 17.1294 |
| Baseline + Q4_K_S | 4-bit | 1.15 | 19.5454 |
| **PreSINQ + Q4_K_S** | 4-bit | 1.01 | **17.4544** |
| Baseline + Q3_K_S | 3-bit | 0.95 | 24.0242 |
| **PreSINQ + Q3_K_S** | 3-bit | 0.83 | **18.8032** |

However, you can generate good PreSINQ models (not the best one) faster by reducing the number of configurations explored during the PreSINQ script execution.
The table below shows perplexity for different PreSINQ parameter configurations using **Q4_K_S quantization**.  
Evaluation is performed on a 5k-line subset of the [**Pile validation dataset**](https://huggingface.co/datasets/mit-han-lab/pile-val-backup).

| Group Size | Iterations | Repetitions | Perplexity |
|-------------|-------------|-------------|-------------|
| 32 | 2 | 1 | 11.7196 |
| 32 | 4 | 1 | 11.7238 |
| 32 | 8 | 1 | **11.6885** |
| 32 | 16 | 1 | 11.6909 |
| 64 | 2 | 1 | 11.7421 |
| 64 | 4 | 1 | 11.7240 |
| 64 | 8 | 1 | 11.6975 |
| 64 | 16 | 1 | 11.7001 |
| 128 | 2 | 1 | 11.7129 |
| 128 | 4 | 1 | 11.7118 |
| 128 | 8 | 1 | 11.7149 |
| 128 | 16 | 1 | 11.7208 |

---

# πŸš€ Usage

## Usage Example

You can load and run the PreSINQ GGUF models using:

- πŸ€— Transformers
- llama.cpp
- Any GGUF-compatible inference framework

---

# 🧾 How to Cite This Work

If you find **SINQ** useful in your research or applications:

- Please give a ⭐ to the official [SINQ](https://github.com/huawei-csl/SINQ) repository  
- Cite our <a href="http://arxiv.org/abs/2509.22944" target="_blank"><strong>paper</strong></a>:

```bibtex
@misc{muller2025sinq,
      title={SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights}, 
      author={Lorenz K. Muller and Philippe Bich and Jiawei Zhuang and Ahmet Celik and Luca Benfenati and Lukas Cavigelli},
      year={2025},
      eprint={2509.22944},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={http://arxiv.org/abs/2509.22944}
}