Card: the project is now TurboQwen (github.com/Ar4ikov/TurboQwen)
Browse files
README.md
CHANGED
|
@@ -24,7 +24,7 @@ tags:
|
|
| 24 |
|
| 25 |
[Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM) prepared for
|
| 26 |
[HyperQwen](https://github.com/syv-ai/HyperQwen) — Qwen3.8-27B served fast on one 24 GB card — the way the
|
| 27 |
-
[
|
| 28 |
AWQ, group 128, zero points, quantized with llm-compressor from bf16 weights; the vision
|
| 29 |
tower, the SSM gate projections and the MTP head's norms stay bf16. What changed is what
|
| 30 |
HyperQwen's `prepare/` pipeline changes, so that a 24 GB card has room for a KV cache and
|
|
@@ -45,7 +45,7 @@ speculative decoding has something small to score:
|
|
| 45 |
The container does everything (download, verify, serve), with the vision tower on:
|
| 46 |
|
| 47 |
```bash
|
| 48 |
-
git clone https://github.com/Ar4ikov/
|
| 49 |
cp .env.example .env # CHECKPOINT=uncensored
|
| 50 |
docker compose --profile single up -d
|
| 51 |
```
|
|
@@ -76,7 +76,7 @@ real prompts with 1,024-token answers; tok/step = tokens accepted per forward pa
|
|
| 76 |
|
| 77 |
Images are described correctly in every profile (a drawn red square, blue circle and a
|
| 78 |
line of text: `boost/image_smoke.py`). The whole table, the int8 (W4A8) profiles and the
|
| 79 |
-
kernel measurements: [github.com/Ar4ikov/
|
| 80 |
|
| 81 |
## Why a separate repo
|
| 82 |
|
|
|
|
| 24 |
|
| 25 |
[Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM) prepared for
|
| 26 |
[HyperQwen](https://github.com/syv-ai/HyperQwen) — Qwen3.8-27B served fast on one 24 GB card — the way the
|
| 27 |
+
[TurboQwen](https://github.com/Ar4ikov/TurboQwen) image expects it. This is the **uncensored** (abliterated) finetune; its behaviour is inherited from [orcarouter/Qwen3.8-27B-Uncensored](https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored). The body is untouched: int4 asymmetric
|
| 28 |
AWQ, group 128, zero points, quantized with llm-compressor from bf16 weights; the vision
|
| 29 |
tower, the SSM gate projections and the MTP head's norms stay bf16. What changed is what
|
| 30 |
HyperQwen's `prepare/` pipeline changes, so that a 24 GB card has room for a KV cache and
|
|
|
|
| 45 |
The container does everything (download, verify, serve), with the vision tower on:
|
| 46 |
|
| 47 |
```bash
|
| 48 |
+
git clone https://github.com/Ar4ikov/TurboQwen && cd TurboQwen
|
| 49 |
cp .env.example .env # CHECKPOINT=uncensored
|
| 50 |
docker compose --profile single up -d
|
| 51 |
```
|
|
|
|
| 76 |
|
| 77 |
Images are described correctly in every profile (a drawn red square, blue circle and a
|
| 78 |
line of text: `boost/image_smoke.py`). The whole table, the int8 (W4A8) profiles and the
|
| 79 |
+
kernel measurements: [github.com/Ar4ikov/TurboQwen](https://github.com/Ar4ikov/TurboQwen).
|
| 80 |
|
| 81 |
## Why a separate repo
|
| 82 |
|