Hakureirm commited on
Commit
95dc8f0
·
verified ·
1 Parent(s): ca75d9d

Add "Where this fits" section; clarify sglang version coupling

Browse files
Files changed (1) hide show
  1. README.md +24 -1
README.md CHANGED
@@ -28,12 +28,35 @@ exchange for needing no calibration data.
28
  - **Speed:** int4 decode faster than fp16 at small batch on every tested arch (identical to
29
  the GPTQ sibling — same kernel).
30
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
31
  ## Format & loading (important)
32
  Not a drop-in HuggingFace checkpoint. Group-wise (GROUP=64) symmetric int4
33
  (`.qweight` + `.scale`); loads **only** through the rwkv-sglang overlay:
34
 
35
  ```bash
36
- bash scripts/deploy.sh # from github.com/Hakureirm/rwkv-sglang, onto sglang v0.5.10.post1
37
  RWKV_W4=1 python -m sglang.launch_server --model-path <this-dir> --dtype float16 \
38
  --trust-remote-code --disable-radix-cache --mem-fraction-static 0.8
39
  ```
 
28
  - **Speed:** int4 decode faster than fp16 at small batch on every tested arch (identical to
29
  the GPTQ sibling — same kernel).
30
 
31
+ ## Where this fits
32
+
33
+ This checkpoint loads through the [rwkv-sglang](https://github.com/Hakureirm/rwkv-sglang)
34
+ overlay. Native RWKV-7 support is being upstreamed into SGLang —
35
+ see [sgl-project/sglang#30115](https://github.com/sgl-project/sglang/pull/30115); once that
36
+ lands the overlay is no longer required.
37
+
38
+ If you want RWKV-7 **without** quantization or a custom runtime, the same base weights are
39
+ also published in standard HuggingFace layout — plain `safetensors`, ordinary `config.json`,
40
+ no `trust_remote_code`:
41
+
42
+ | | |
43
+ |---|---|
44
+ | 0.1B | [rwkv7-0.1b-hf](https://huggingface.co/Hakureirm/rwkv7-0.1b-hf) |
45
+ | 0.4B | [rwkv7-0.4b-hf](https://huggingface.co/Hakureirm/rwkv7-0.4b-hf) |
46
+ | 1.5B | [rwkv7-1.5b-hf](https://huggingface.co/Hakureirm/rwkv7-1.5b-hf) |
47
+ | 2.9B | [rwkv7-2.9b-hf](https://huggingface.co/Hakureirm/rwkv7-2.9b-hf) |
48
+ | 7.2B (G1) | [rwkv7-g1h-7.2b-hf](https://huggingface.co/Hakureirm/rwkv7-g1h-7.2b-hf) |
49
+ | Pile 168M | [rwkv7-168m-pile-hf](https://huggingface.co/Hakureirm/rwkv7-168m-pile-hf) |
50
+
51
+ Transformers support for the architecture itself is open as
52
+ [huggingface/transformers#47780](https://github.com/huggingface/transformers/pull/47780).
53
+
54
  ## Format & loading (important)
55
  Not a drop-in HuggingFace checkpoint. Group-wise (GROUP=64) symmetric int4
56
  (`.qweight` + `.scale`); loads **only** through the rwkv-sglang overlay:
57
 
58
  ```bash
59
+ bash scripts/deploy.sh # from github.com/Hakureirm/rwkv-sglang, built against sglang v0.5.10.post1 — newer releases need the overlay rebased
60
  RWKV_W4=1 python -m sglang.launch_server --model-path <this-dir> --dtype float16 \
61
  --trust-remote-code --disable-radix-cache --mem-fraction-static 0.8
62
  ```