File size: 926 Bytes
2c36cb3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 | ---
license: apache-2.0
base_model: Qwen/Qwen3-4B-Thinking-2507
tags:
- quantized
- rtn
- 4-bit
- thinking
- reasoning
---
# Qwen__Qwen3-4B-Thinking-2507_RTN_w4g128
This is a 4-bit RTN (Round-To-Nearest) quantized version of [Qwen/Qwen3-4B-Thinking-2507](https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507).
## Quantization Details
- **Method**: RTN (Round-To-Nearest)
- **Bits**: 4-bit
- **Group Size**: 128
- **Base Model**: [Qwen/Qwen3-4B-Thinking-2507](https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507)
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "quantpa/Qwen__Qwen3-4B-Thinking-2507_RTN_w4g128"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
# Use the model for inference
```
## Model Details
- **Quantization**: RTN 4-bit
- **Original Model**: Qwen/Qwen3-4B-Thinking-2507
- **Quantized by**: quantpa
|