File size: 926 Bytes
2c36cb3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
---
license: apache-2.0
base_model: Qwen/Qwen3-4B-Thinking-2507
tags:
- quantized
- rtn
- 4-bit
- thinking
- reasoning
---

# Qwen__Qwen3-4B-Thinking-2507_RTN_w4g128

This is a 4-bit RTN (Round-To-Nearest) quantized version of [Qwen/Qwen3-4B-Thinking-2507](https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507).

## Quantization Details

- **Method**: RTN (Round-To-Nearest)
- **Bits**: 4-bit
- **Group Size**: 128
- **Base Model**: [Qwen/Qwen3-4B-Thinking-2507](https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507)

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "quantpa/Qwen__Qwen3-4B-Thinking-2507_RTN_w4g128"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

# Use the model for inference
```

## Model Details

- **Quantization**: RTN 4-bit
- **Original Model**: Qwen/Qwen3-4B-Thinking-2507
- **Quantized by**: quantpa