Ar4ikov commited on
Commit
9bec9de
·
verified ·
1 Parent(s): d3c097b

Card: the project is now TurboQwen (github.com/Ar4ikov/TurboQwen)

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -24,7 +24,7 @@ tags:
24
 
25
  [Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM) prepared for
26
  [HyperQwen](https://github.com/syv-ai/HyperQwen) — Qwen3.8-27B served fast on one 24 GB card — the way the
27
- [vllm-hyprfastQwen](https://github.com/Ar4ikov/vllm-hyprfastQwen) image expects it. This is the **uncensored** (abliterated) finetune; its behaviour is inherited from [orcarouter/Qwen3.8-27B-Uncensored](https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored). The body is untouched: int4 asymmetric
28
  AWQ, group 128, zero points, quantized with llm-compressor from bf16 weights; the vision
29
  tower, the SSM gate projections and the MTP head's norms stay bf16. What changed is what
30
  HyperQwen's `prepare/` pipeline changes, so that a 24 GB card has room for a KV cache and
@@ -45,7 +45,7 @@ speculative decoding has something small to score:
45
  The container does everything (download, verify, serve), with the vision tower on:
46
 
47
  ```bash
48
- git clone https://github.com/Ar4ikov/vllm-hyprfastQwen && cd vllm-hyprfastQwen
49
  cp .env.example .env # CHECKPOINT=uncensored
50
  docker compose --profile single up -d
51
  ```
@@ -76,7 +76,7 @@ real prompts with 1,024-token answers; tok/step = tokens accepted per forward pa
76
 
77
  Images are described correctly in every profile (a drawn red square, blue circle and a
78
  line of text: `boost/image_smoke.py`). The whole table, the int8 (W4A8) profiles and the
79
- kernel measurements: [github.com/Ar4ikov/vllm-hyprfastQwen](https://github.com/Ar4ikov/vllm-hyprfastQwen).
80
 
81
  ## Why a separate repo
82
 
 
24
 
25
  [Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM](https://huggingface.co/Ar4ikov/Qwen3.8-27B-Uncensored-AWQ-W4A16-ASYM) prepared for
26
  [HyperQwen](https://github.com/syv-ai/HyperQwen) — Qwen3.8-27B served fast on one 24 GB card — the way the
27
+ [TurboQwen](https://github.com/Ar4ikov/TurboQwen) image expects it. This is the **uncensored** (abliterated) finetune; its behaviour is inherited from [orcarouter/Qwen3.8-27B-Uncensored](https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored). The body is untouched: int4 asymmetric
28
  AWQ, group 128, zero points, quantized with llm-compressor from bf16 weights; the vision
29
  tower, the SSM gate projections and the MTP head's norms stay bf16. What changed is what
30
  HyperQwen's `prepare/` pipeline changes, so that a 24 GB card has room for a KV cache and
 
45
  The container does everything (download, verify, serve), with the vision tower on:
46
 
47
  ```bash
48
+ git clone https://github.com/Ar4ikov/TurboQwen && cd TurboQwen
49
  cp .env.example .env # CHECKPOINT=uncensored
50
  docker compose --profile single up -d
51
  ```
 
76
 
77
  Images are described correctly in every profile (a drawn red square, blue circle and a
78
  line of text: `boost/image_smoke.py`). The whole table, the int8 (W4A8) profiles and the
79
+ kernel measurements: [github.com/Ar4ikov/TurboQwen](https://github.com/Ar4ikov/TurboQwen).
80
 
81
  ## Why a separate repo
82