Instructions to use Hakureirm/rwkv7-sglang-w4rtn-7.2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RWKV
How to use Hakureirm/rwkv7-sglang-w4rtn-7.2b with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Add "Where this fits" section; clarify sglang version coupling
Browse files
README.md
CHANGED
|
@@ -28,12 +28,35 @@ exchange for needing no calibration data.
|
|
| 28 |
- **Speed:** int4 decode faster than fp16 at small batch on every tested arch (identical to
|
| 29 |
the GPTQ sibling — same kernel).
|
| 30 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 31 |
## Format & loading (important)
|
| 32 |
Not a drop-in HuggingFace checkpoint. Group-wise (GROUP=64) symmetric int4
|
| 33 |
(`.qweight` + `.scale`); loads **only** through the rwkv-sglang overlay:
|
| 34 |
|
| 35 |
```bash
|
| 36 |
-
bash scripts/deploy.sh # from github.com/Hakureirm/rwkv-sglang,
|
| 37 |
RWKV_W4=1 python -m sglang.launch_server --model-path <this-dir> --dtype float16 \
|
| 38 |
--trust-remote-code --disable-radix-cache --mem-fraction-static 0.8
|
| 39 |
```
|
|
|
|
| 28 |
- **Speed:** int4 decode faster than fp16 at small batch on every tested arch (identical to
|
| 29 |
the GPTQ sibling — same kernel).
|
| 30 |
|
| 31 |
+
## Where this fits
|
| 32 |
+
|
| 33 |
+
This checkpoint loads through the [rwkv-sglang](https://github.com/Hakureirm/rwkv-sglang)
|
| 34 |
+
overlay. Native RWKV-7 support is being upstreamed into SGLang —
|
| 35 |
+
see [sgl-project/sglang#30115](https://github.com/sgl-project/sglang/pull/30115); once that
|
| 36 |
+
lands the overlay is no longer required.
|
| 37 |
+
|
| 38 |
+
If you want RWKV-7 **without** quantization or a custom runtime, the same base weights are
|
| 39 |
+
also published in standard HuggingFace layout — plain `safetensors`, ordinary `config.json`,
|
| 40 |
+
no `trust_remote_code`:
|
| 41 |
+
|
| 42 |
+
| | |
|
| 43 |
+
|---|---|
|
| 44 |
+
| 0.1B | [rwkv7-0.1b-hf](https://huggingface.co/Hakureirm/rwkv7-0.1b-hf) |
|
| 45 |
+
| 0.4B | [rwkv7-0.4b-hf](https://huggingface.co/Hakureirm/rwkv7-0.4b-hf) |
|
| 46 |
+
| 1.5B | [rwkv7-1.5b-hf](https://huggingface.co/Hakureirm/rwkv7-1.5b-hf) |
|
| 47 |
+
| 2.9B | [rwkv7-2.9b-hf](https://huggingface.co/Hakureirm/rwkv7-2.9b-hf) |
|
| 48 |
+
| 7.2B (G1) | [rwkv7-g1h-7.2b-hf](https://huggingface.co/Hakureirm/rwkv7-g1h-7.2b-hf) |
|
| 49 |
+
| Pile 168M | [rwkv7-168m-pile-hf](https://huggingface.co/Hakureirm/rwkv7-168m-pile-hf) |
|
| 50 |
+
|
| 51 |
+
Transformers support for the architecture itself is open as
|
| 52 |
+
[huggingface/transformers#47780](https://github.com/huggingface/transformers/pull/47780).
|
| 53 |
+
|
| 54 |
## Format & loading (important)
|
| 55 |
Not a drop-in HuggingFace checkpoint. Group-wise (GROUP=64) symmetric int4
|
| 56 |
(`.qweight` + `.scale`); loads **only** through the rwkv-sglang overlay:
|
| 57 |
|
| 58 |
```bash
|
| 59 |
+
bash scripts/deploy.sh # from github.com/Hakureirm/rwkv-sglang, built against sglang v0.5.10.post1 — newer releases need the overlay rebased
|
| 60 |
RWKV_W4=1 python -m sglang.launch_server --model-path <this-dir> --dtype float16 \
|
| 61 |
--trust-remote-code --disable-radix-cache --mem-fraction-static 0.8
|
| 62 |
```
|