shoumenchougou commited on
Commit
3a89d27
·
verified ·
1 Parent(s): 7d03bc8

Publish verified RWKV7 G1k 2.9B release

Browse files
LICENSE ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and do not
117
+ modify the License. You may add Your own attribution notices
118
+ within Derivative Works that You distribute, alongside or as an
119
+ addendum to the NOTICE text from the Work, provided that such
120
+ additional attribution notices cannot be construed as modifying
121
+ the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer, and
167
+ charge a fee for, acceptance of support, warranty, indemnity, or
168
+ other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. We also recommend that a
185
+ file or class name and description of purpose be included on the
186
+ same "printed page" as the copyright notice for easier
187
+ identification within third-party archives.
188
+
189
+ Copyright [yyyy] [name of copyright owner]
190
+
191
+ Licensed under the Apache License, Version 2.0 (the "License");
192
+ you may not use this file except in compliance with the License.
193
+ You may obtain a copy of the License at
194
+
195
+ http://www.apache.org/licenses/LICENSE-2.0
196
+
197
+ Unless required by applicable law or agreed to in writing, software
198
+ distributed under the License is distributed on an "AS IS" BASIS,
199
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
200
+ See the License for the specific language governing permissions and
201
+ limitations under the License.
NOTICE ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ RWKV-7 G1k model release
2
+
3
+ Source checkpoint: BlinkDL/rwkv7-g1/rwkv7-g1k-2.9b-20260930-ctx25600.pth
4
+ Source revision: a1ddf9e96df23d5d7db65137ad83a4701044cd05
5
+ Source SHA-256: d8f8ecd4af8d706f867ac8a74284f35c0e3b9f863e7b09bb4d4031480c8b28b2
6
+ FLA implementation revision: 8e84ed4a6727be082c34a3855c60623fd11411e9
7
+
8
+ The RWKV logo in assets/rwkv-logo.webp is the audited Hugging Face RWKV organization avatar snapshot. RWKV names and logos may be subject to separate trademark rules; Apache-2.0 does not grant trademark rights.
README.md ADDED
@@ -0,0 +1,153 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ pipeline_tag: text-generation
5
+ base_model: BlinkDL/rwkv7-g1
6
+ language:
7
+ - en
8
+ - zh
9
+ - fr
10
+ - es
11
+ - de
12
+ - pt
13
+ - ru
14
+ - it
15
+ - ja
16
+ - ko
17
+ - vi
18
+ - ar
19
+ datasets:
20
+ - HuggingFaceFW/fineweb-edu
21
+ - mlfoundations/dclm-baseline-1.0
22
+ - cerebras/SlimPajama-627B
23
+ - EleutherAI/pile
24
+ - bigcode/starcoderdata
25
+ - oscar-corpus/OSCAR-2301
26
+ tags:
27
+ - rwkv
28
+ - rwkv7
29
+ - fla
30
+ - recurrent
31
+ - causal-lm
32
+ ---
33
+
34
+ <!-- markdownlint-disable first-line-h1 -->
35
+ <!-- markdownlint-disable html -->
36
+
37
+ <div align="center">
38
+ <a href="https://www.rwkv.com/">
39
+ <img src="assets/rwkv-logo.webp" width="140" alt="RWKV logo" />
40
+ </a>
41
+ <h1>RWKV7-G1k-2.9B-20260930</h1>
42
+ <p><strong>RWKV-7 “Goose” in Flash Linear Attention format</strong></p>
43
+ </div>
44
+
45
+ This repository provides the **RWKV-7 G1k 2.9B** checkpoint in the
46
+ [`flash-linear-attention`](https://github.com/fla-org/flash-linear-attention)
47
+ (FLA) RWKV7 layout for Transformers-compatible inference.
48
+
49
+ > [!IMPORTANT]
50
+ > This is a base language model, not a safety-aligned instruction-tuned
51
+ > assistant. The included chat template provides a conversational prompt
52
+ > format, but the model may not follow instructions consistently.
53
+
54
+ ## About RWKV-7
55
+
56
+ RWKV-7, also called **Goose**, is an attention-free recurrent language model.
57
+ It maintains a constant-size recurrent state instead of an attention KV cache
58
+ that grows with the preceding sequence. Training remains parallelizable, while
59
+ recurrent decoding uses constant state size and constant work per generated
60
+ token with respect to sequence length.
61
+
62
+ `G1k` identifies this checkpoint revision. The source checkpoint is available
63
+ from [`BlinkDL/rwkv7-g1`](https://huggingface.co/BlinkDL/rwkv7-g1).
64
+
65
+ ## Model details
66
+
67
+ | Property | Value |
68
+ | --- | --- |
69
+ | Architecture | RWKV-7 G1k |
70
+ | Parameters in this FLA checkpoint | 2.948B |
71
+ | Layers | 32 |
72
+ | Hidden size | 2,560 |
73
+ | Heads | 40 × 64 |
74
+ | Feed-forward size | 10,240 |
75
+ | Vocabulary | RWKV World tokenizer, 65,536 tokens |
76
+ | Configured context length | 25,600 tokens |
77
+ | Weight dtype | BF16 |
78
+ | License | Apache-2.0 |
79
+
80
+ ## Run with FLA
81
+
82
+ Use a CUDA-capable NVIDIA GPU with BF16 support. Install the audited FLA
83
+ revision with its CUDA dependency extra:
84
+
85
+ ```bash
86
+ python -m pip install \
87
+ "flash-linear-attention[cuda] @ git+https://github.com/fla-org/flash-linear-attention.git@8e84ed4a6727be082c34a3855c60623fd11411e9" \
88
+ "transformers>=4.50.2"
89
+ ```
90
+
91
+ Import `fla` before using the Auto classes so that the RWKV7 implementation is
92
+ registered:
93
+
94
+ ```python
95
+ import fla
96
+ import torch
97
+ from transformers import AutoModelForCausalLM, AutoTokenizer, PretrainedConfig
98
+
99
+ model_id = "shoumenchougou/RWKV7-G1k-2.9B-20260930"
100
+
101
+ tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True, config=PretrainedConfig())
102
+ model = AutoModelForCausalLM.from_pretrained(
103
+ model_id,
104
+ torch_dtype=torch.bfloat16,
105
+ ).to("cuda").eval()
106
+
107
+ messages = [
108
+ {"role": "user", "content": "Explain recurrent language models in one paragraph."}
109
+ ]
110
+ prompt = tokenizer.apply_chat_template(
111
+ messages,
112
+ tokenize=False,
113
+ add_generation_prompt=True,
114
+ thinking=False,
115
+ )
116
+ inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
117
+
118
+ with torch.inference_mode():
119
+ output_ids = model.generate(
120
+ **inputs,
121
+ max_new_tokens=128,
122
+ do_sample=True,
123
+ temperature=0.7,
124
+ top_p=0.9,
125
+ )
126
+
127
+ new_tokens = output_ids[0, inputs.input_ids.shape[1]:]
128
+ print(tokenizer.decode(new_tokens, skip_special_tokens=True))
129
+ ```
130
+
131
+ Set `thinking=True` when applying the chat template to leave an open thinking
132
+ prefix. This changes only the prompt format; it does not turn the base model
133
+ into an instruction-tuned or safety-aligned assistant.
134
+
135
+ ## Compatibility and validation
136
+
137
+ The package was checked with FLA commit
138
+ `8e84ed4a6727be082c34a3855c60623fd11411e9` and Transformers 4.50.2. Its
139
+ configuration, tokenizer, chat template, BF16 weights, and Transformers model
140
+ loading were validated locally. All 1,059 model tensors load into
141
+ `RWKV7ForCausalLM` without missing, unexpected, or mismatched keys.
142
+
143
+ CUDA/Triton generation was not executed in the local CPU-only validation
144
+ environment. Runtime behavior can depend on the GPU, CUDA, PyTorch, Triton,
145
+ and FLA versions. Cross-check benchmark or evaluation results against the
146
+ official RWKV implementation before reporting them.
147
+
148
+ ## References
149
+
150
+ - [RWKV-7 paper](https://arxiv.org/abs/2503.14456)
151
+ - [RWKV-LM](https://github.com/BlinkDL/RWKV-LM)
152
+ - [Flash Linear Attention](https://github.com/fla-org/flash-linear-attention)
153
+ - [Source checkpoint repository](https://huggingface.co/BlinkDL/rwkv7-g1)
assets/rwkv-logo.webp ADDED
chat_template.jinja ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set add_generation_prompt = add_generation_prompt | default(false) -%}
2
+ {%- set thinking = thinking | default(false) -%}
3
+ {%- set bos_token = bos_token | default('', true) -%}
4
+ {%- set tools = tools | default([], true) -%}
5
+ {%- set ns = namespace(system_prompt='') -%}
6
+ {%- for message in messages -%}
7
+ {%- if message.role == 'system' -%}
8
+ {%- set ns.system_prompt = message.content | trim -%}
9
+ {%- endif -%}
10
+ {%- endfor -%}
11
+ {{- bos_token -}}
12
+ {%- if ns.system_prompt or tools | length > 0 -%}
13
+ {{ 'System: ' }}{{ ns.system_prompt }}
14
+ {%- if tools | length > 0 -%}
15
+ {%- if ns.system_prompt %}{{ '\n' }}{%- endif -%}
16
+ {{ 'Tools:\n' -}}
17
+ {{ tools | tojson }}
18
+ {{ '\nWhen using a tool, return only a compact JSON function call in a ```json block, like {"name":"calculator","arguments":{"expression":"2+2"}}. The `name` field must be top-level, never inside `arguments`. Do not copy the tool schema into arguments. Otherwise answer normally.' }}
19
+ {%- endif -%}
20
+ {{ '\n\n' }}
21
+ {%- endif -%}
22
+ {%- for message in messages -%}
23
+ {%- if message.role == 'user' -%}
24
+ {{ 'User: ' ~ (message.content | trim) ~ '\n\n' }}
25
+ {%- elif message.role == 'assistant' -%}
26
+ {%- generation -%}
27
+ {%- set content = message.content | default('', true) | trim -%}
28
+ {{ 'Assistant:' }}
29
+ {%- if message.tool_calls is defined and message.tool_calls | length > 0 -%}
30
+ {%- if content %}{{ ' ' ~ content ~ '\n' }}{%- endif -%}
31
+ {%- for tool_call in message.tool_calls -%}
32
+ {%- if tool_call.function is defined -%}
33
+ {%- set name = tool_call.function.name | default('') -%}
34
+ {%- set args = tool_call.function.arguments | default({}, true) -%}
35
+ {%- else -%}
36
+ {%- set name = tool_call.name | default('') -%}
37
+ {%- set args = tool_call.arguments | default({}, true) -%}
38
+ {%- endif -%}
39
+ {{ ' ```json\n' -}}
40
+ {{ '{"name": ' }}{{ name | tojson }}{{ ', "arguments": ' }}{% if args is string %}{{ args }}{% else %}{{ args | tojson }}{% endif %}{{ '}\n' -}}
41
+ {{ '```' }}{{ '\n' if not loop.last else '' }}
42
+ {%- endfor -%}
43
+ {%- elif content -%}
44
+ {{ ' ' ~ content }}
45
+ {%- endif -%}
46
+ {{ '\n\n' }}
47
+ {%- endgeneration -%}
48
+ {%- elif message.role == 'tool' -%}
49
+ {{ 'User: Function output:\n' ~ (message.content | trim) ~ '\n\n' }}
50
+ {%- endif -%}
51
+ {%- endfor -%}
52
+ {%- if add_generation_prompt -%}
53
+ {%- if thinking -%}
54
+ {{ 'Assistant: <think' }}
55
+ {%- else -%}
56
+ {{ 'Assistant: <think></think>\n' }}
57
+ {%- endif -%}
58
+ {%- endif -%}
config.json ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "a_low_rank_dim": 96,
3
+ "architectures": [
4
+ "RWKV7ForCausalLM"
5
+ ],
6
+ "attn": null,
7
+ "attn_mode": "chunk",
8
+ "bos_token_id": 0,
9
+ "decay_low_rank_dim": 96,
10
+ "eos_token_id": 0,
11
+ "fuse_cross_entropy": true,
12
+ "fuse_linear_cross_entropy": false,
13
+ "fuse_norm": true,
14
+ "gate_low_rank_dim": 320,
15
+ "head_dim": 64,
16
+ "hidden_act": "sqrelu",
17
+ "hidden_ratio": 4.0,
18
+ "hidden_size": 2560,
19
+ "initializer_range": 0.02,
20
+ "intermediate_size": 10240,
21
+ "max_position_embeddings": 25600,
22
+ "model_type": "rwkv7",
23
+ "norm_bias": true,
24
+ "norm_eps": 1e-05,
25
+ "norm_first": true,
26
+ "num_heads": 40,
27
+ "num_hidden_layers": 32,
28
+ "tie_word_embeddings": false,
29
+ "torch_dtype": "bfloat16",
30
+ "transformers_version": "4.50.2",
31
+ "use_cache": true,
32
+ "use_l2warp": true,
33
+ "v_low_rank_dim": 64,
34
+ "value_dim": [
35
+ 2560,
36
+ 2560,
37
+ 2560,
38
+ 2560,
39
+ 2560,
40
+ 2560,
41
+ 2560,
42
+ 2560,
43
+ 2560,
44
+ 2560,
45
+ 2560,
46
+ 2560,
47
+ 2560,
48
+ 2560,
49
+ 2560,
50
+ 2560,
51
+ 2560,
52
+ 2560,
53
+ 2560,
54
+ 2560,
55
+ 2560,
56
+ 2560,
57
+ 2560,
58
+ 2560,
59
+ 2560,
60
+ 2560,
61
+ 2560,
62
+ 2560,
63
+ 2560,
64
+ 2560,
65
+ 2560,
66
+ 2560
67
+ ],
68
+ "vocab_size": 65536,
69
+ "pad_token_id": 0
70
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 0,
3
+ "eos_token_id": 0,
4
+ "pad_token_id": 0,
5
+ "use_cache": true
6
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:92b14263e2c78aa2a07b636e23395726f15ce3b53e2a5cfaea6681ec7215f4fb
3
+ size 5895584920
tokenization_rwkv7.py ADDED
@@ -0,0 +1,343 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # coding=utf-8
2
+ """Exact linear-time RWKV World tokenizer for Transformers AutoTokenizer."""
3
+
4
+ from __future__ import annotations
5
+
6
+ import json
7
+ import os
8
+ import shutil
9
+ from typing import Optional
10
+
11
+ from transformers.tokenization_utils import PreTrainedTokenizer
12
+
13
+
14
+ VOCAB_FILES_NAMES = {"tokenizer_file": "tokenizer.json"}
15
+ END_TOKEN = "<|endoftext|>"
16
+ INVALID_TOKEN_PREFIX = "\ue000"
17
+ INVALID_TOKEN_SUFFIX = "\ue001"
18
+
19
+
20
+ class _CharEncoding:
21
+ """Minimal Encoding surface used by BatchEncoding.char_to_token."""
22
+
23
+ def __init__(self, offsets):
24
+ self.offsets = offsets
25
+
26
+ def char_to_token(self, char_index, sequence_index=0):
27
+ if sequence_index != 0:
28
+ return None
29
+ for token_index, offset in enumerate(self.offsets):
30
+ if offset is not None and offset[0] <= char_index < offset[1]:
31
+ return token_index
32
+ return None
33
+
34
+
35
+ def _bytes_to_unicode():
36
+ values = list(range(ord("!"), ord("~") + 1))
37
+ values += list(range(ord("¡"), ord("¬") + 1))
38
+ values += list(range(ord("®"), ord("ÿ") + 1))
39
+ characters = list(values)
40
+ extra = 0
41
+ for byte in range(256):
42
+ if byte not in values:
43
+ values.append(byte)
44
+ characters.append(256 + extra)
45
+ extra += 1
46
+ return dict(zip(values, map(chr, characters)))
47
+
48
+
49
+ class Rwkv7Tokenizer(PreTrainedTokenizer):
50
+ """RWKV World longest-prefix tokenizer backed only by tokenizer.json."""
51
+
52
+ vocab_files_names = VOCAB_FILES_NAMES
53
+ model_input_names = ["input_ids", "attention_mask"]
54
+ padding_side = "left"
55
+
56
+ def __init__(
57
+ self,
58
+ tokenizer_file,
59
+ eos_token=END_TOKEN,
60
+ pad_token=END_TOKEN,
61
+ unk_token=END_TOKEN,
62
+ add_bos_token=False,
63
+ **kwargs,
64
+ ):
65
+ if not tokenizer_file or not os.path.isfile(tokenizer_file):
66
+ raise ValueError(f"tokenizer.json does not exist: {tokenizer_file}")
67
+ if add_bos_token:
68
+ raise ValueError("RWKV World has no BOS token")
69
+ self.tokenizer_file = tokenizer_file
70
+ self.add_bos_token = False
71
+ value = json.loads(open(tokenizer_file, encoding="utf-8").read())
72
+ model = value.get("model")
73
+ if not isinstance(model, dict) or model.get("type") != "WordPiece":
74
+ raise ValueError("tokenizer.json is not the locked RWKV WordPiece model")
75
+ if (
76
+ model.get("unk_token") != END_TOKEN
77
+ or model.get("continuing_subword_prefix") != ""
78
+ or value.get("normalizer") is not None
79
+ ):
80
+ raise ValueError("tokenizer.json does not use exact RWKV semantics")
81
+ pre_tokenizer = value.get("pre_tokenizer")
82
+ decoder = value.get("decoder")
83
+ if (
84
+ not isinstance(pre_tokenizer, dict)
85
+ or pre_tokenizer.get("type") != "ByteLevel"
86
+ or pre_tokenizer.get("add_prefix_space") is not False
87
+ or pre_tokenizer.get("use_regex") is not False
88
+ or not isinstance(decoder, dict)
89
+ or decoder.get("type") != "ByteLevel"
90
+ ):
91
+ raise ValueError("tokenizer.json does not use exact ByteLevel semantics")
92
+ vocabulary = model.get("vocab")
93
+ if not isinstance(vocabulary, dict) or not vocabulary:
94
+ raise ValueError("tokenizer.json has no vocabulary")
95
+ if any(not isinstance(token, str) or not isinstance(index, int) for token, index in vocabulary.items()):
96
+ raise TypeError("tokenizer.json vocabulary entry is invalid")
97
+ if set(vocabulary.values()) != set(range(len(vocabulary))):
98
+ raise ValueError("tokenizer.json vocabulary ids are not contiguous")
99
+ if vocabulary.get(END_TOKEN) != 0:
100
+ raise ValueError("tokenizer.json must reserve id 0 for end-of-text")
101
+
102
+ byte_decoder = {character: byte for byte, character in _bytes_to_unicode().items()}
103
+ self.encoder = dict(vocabulary)
104
+ self.decoder = {index: token for token, index in vocabulary.items()}
105
+ self._token_bytes = {}
106
+ reverse = {}
107
+ for token, index in vocabulary.items():
108
+ if index == 0:
109
+ continue
110
+ placeholder = f"{INVALID_TOKEN_PREFIX}{index}{INVALID_TOKEN_SUFFIX}"
111
+ if token == placeholder:
112
+ continue
113
+ try:
114
+ raw = bytes(byte_decoder[character] for character in token)
115
+ except KeyError as error:
116
+ raise ValueError("tokenizer.json vocabulary is not ByteLevel encoded") from error
117
+ if raw in reverse:
118
+ raise ValueError("tokenizer.json contains duplicate byte tokens")
119
+ reverse[raw] = index
120
+ self._token_bytes[token] = raw
121
+ missing = [byte for byte in range(256) if bytes([byte]) not in reverse]
122
+ if missing:
123
+ raise ValueError(f"tokenizer.json is missing singleton bytes: {missing}")
124
+ for raw, index in reverse.items():
125
+ for endpoint in range(1, len(raw)):
126
+ prefix = reverse.get(raw[:endpoint])
127
+ if prefix is not None and prefix > index:
128
+ raise ValueError("token ranks do not support longest-prefix encoding")
129
+
130
+ self._children = [{}]
131
+ self._terminal = [None]
132
+ self._max_token_bytes = 1
133
+ for raw, index in sorted(reverse.items(), key=lambda item: item[1]):
134
+ node = 0
135
+ for byte in raw:
136
+ child = self._children[node].get(byte)
137
+ if child is None:
138
+ child = len(self._children)
139
+ self._children[node][byte] = child
140
+ self._children.append({})
141
+ self._terminal.append(None)
142
+ node = child
143
+ self._terminal[node] = index
144
+ self._max_token_bytes = max(self._max_token_bytes, len(raw))
145
+
146
+ if "additional_special_tokens" not in kwargs:
147
+ appended = []
148
+ for item in sorted(value.get("added_tokens", []), key=lambda item: item.get("id", -1)):
149
+ if not isinstance(item, dict):
150
+ raise TypeError("tokenizer.json added token is invalid")
151
+ content = item.get("content")
152
+ index = item.get("id")
153
+ if not isinstance(content, str) or not isinstance(index, int):
154
+ raise TypeError("tokenizer.json added token is invalid")
155
+ if content != END_TOKEN:
156
+ if not item.get("special") or index < len(vocabulary):
157
+ raise ValueError("only append-only special tokens are supported")
158
+ appended.append(content)
159
+ if appended:
160
+ kwargs["additional_special_tokens"] = appended
161
+ super().__init__(
162
+ eos_token=eos_token,
163
+ pad_token=pad_token,
164
+ unk_token=unk_token,
165
+ add_bos_token=self.add_bos_token,
166
+ **kwargs,
167
+ )
168
+
169
+ @property
170
+ def vocab_size(self):
171
+ return len(self.encoder)
172
+
173
+ def get_vocab(self):
174
+ vocabulary = dict(self.encoder)
175
+ vocabulary.update(self.added_tokens_encoder)
176
+ return vocabulary
177
+
178
+ def _encode_bytes(self, data):
179
+ output = []
180
+ position = 0
181
+ while position < len(data):
182
+ node = 0
183
+ cursor = position
184
+ best_id = None
185
+ best_end = position
186
+ limit = min(len(data), position + self._max_token_bytes)
187
+ while cursor < limit:
188
+ child = self._children[node].get(data[cursor])
189
+ if child is None:
190
+ break
191
+ node = child
192
+ cursor += 1
193
+ token_id = self._terminal[node]
194
+ if token_id is not None and (best_id is None or token_id > best_id):
195
+ best_id = token_id
196
+ best_end = cursor
197
+ if best_id is None:
198
+ raise RuntimeError("singleton-byte vocabulary invariant was violated")
199
+ output.append(best_id)
200
+ position = best_end
201
+ return output
202
+
203
+ def _encode_bytes_with_offsets(self, data, char_offset, byte_to_char):
204
+ output = []
205
+ position = 0
206
+ while position < len(data):
207
+ node = 0
208
+ cursor = position
209
+ best_id = None
210
+ best_end = position
211
+ limit = min(len(data), position + self._max_token_bytes)
212
+ while cursor < limit:
213
+ child = self._children[node].get(data[cursor])
214
+ if child is None:
215
+ break
216
+ node = child
217
+ cursor += 1
218
+ token_id = self._terminal[node]
219
+ if token_id is not None and (best_id is None or token_id > best_id):
220
+ best_id = token_id
221
+ best_end = cursor
222
+ if best_id is None:
223
+ raise RuntimeError("singleton-byte vocabulary invariant was violated")
224
+ start_char = char_offset + byte_to_char[position]
225
+ end_char = char_offset + byte_to_char[best_end - 1] + 1
226
+ output.append((best_id, (start_char, end_char)))
227
+ position = best_end
228
+ return output
229
+
230
+ def _encode_with_offsets(self, text):
231
+ special = sorted(self.added_tokens_encoder, key=lambda token: (-len(token), token))
232
+ output = []
233
+ ordinary_start = 0
234
+ position = 0
235
+
236
+ def append_ordinary(segment, char_offset):
237
+ if not segment:
238
+ return
239
+ data = segment.encode("utf-8")
240
+ byte_to_char = []
241
+ for char_index, character in enumerate(segment):
242
+ byte_to_char.extend([char_index] * len(character.encode("utf-8")))
243
+ output.extend(
244
+ self._encode_bytes_with_offsets(data, char_offset, byte_to_char)
245
+ )
246
+
247
+ while position < len(text):
248
+ matched = next(
249
+ (token for token in special if text.startswith(token, position)), None
250
+ )
251
+ if matched is None:
252
+ position += 1
253
+ continue
254
+ append_ordinary(text[ordinary_start:position], ordinary_start)
255
+ output.append(
256
+ (
257
+ self.added_tokens_encoder[matched],
258
+ (position, position + len(matched)),
259
+ )
260
+ )
261
+ position += len(matched)
262
+ ordinary_start = position
263
+ append_ordinary(text[ordinary_start:], ordinary_start)
264
+ return output
265
+
266
+ def _tokenize(self, text, **kwargs):
267
+ del kwargs
268
+ return [self.decoder[index] for index in self._encode_bytes(text.encode("utf-8"))]
269
+
270
+ def __call__(self, text=None, text_pair=None, **kwargs):
271
+ output = super().__call__(text=text, text_pair=text_pair, **kwargs)
272
+ if text_pair is not None or text is None:
273
+ return output
274
+ texts = [text] if isinstance(text, str) else list(text)
275
+ if any(not isinstance(item, str) for item in texts):
276
+ return output
277
+ input_ids = output["input_ids"]
278
+ attention_mask = output.get("attention_mask")
279
+ raw_ids = input_ids.tolist() if hasattr(input_ids, "tolist") else input_ids
280
+ if raw_ids and isinstance(raw_ids[0], int):
281
+ rows = [raw_ids]
282
+ masks = [attention_mask] if attention_mask is not None else None
283
+ else:
284
+ rows = raw_ids
285
+ if attention_mask is None:
286
+ masks = None
287
+ else:
288
+ masks = (
289
+ attention_mask.tolist()
290
+ if hasattr(attention_mask, "tolist")
291
+ else attention_mask
292
+ )
293
+ encodings = []
294
+ for index, source in enumerate(texts):
295
+ exact = self._encode_with_offsets(source)
296
+ row = rows[index]
297
+ active = len(row) if masks is None else sum(int(item) for item in masks[index])
298
+ exact = exact[:active]
299
+ offsets = [offset for _token_id, offset in exact]
300
+ missing = len(row) - len(offsets)
301
+ if self.padding_side == "left":
302
+ offsets = [None] * missing + offsets
303
+ else:
304
+ offsets.extend([None] * missing)
305
+ encodings.append(_CharEncoding(offsets))
306
+ output._encodings = encodings
307
+ return output
308
+
309
+ def _convert_token_to_id(self, token):
310
+ return self.encoder.get(token, 0)
311
+
312
+ def _convert_id_to_token(self, index):
313
+ return self.decoder.get(index, END_TOKEN)
314
+
315
+ def convert_tokens_to_string(self, tokens):
316
+ decoded = bytearray()
317
+ for token in tokens:
318
+ raw = self._token_bytes.get(token)
319
+ if raw is not None:
320
+ decoded.extend(raw)
321
+ else:
322
+ decoded.extend(str(token).encode("utf-8"))
323
+ return bytes(decoded).decode("utf-8", errors="replace")
324
+
325
+ def build_inputs_with_special_tokens(self, token_ids_0, token_ids_1=None):
326
+ bos = [self.bos_token_id] if self.add_bos_token else []
327
+ output = bos + list(token_ids_0)
328
+ if token_ids_1 is not None:
329
+ output.extend(bos + list(token_ids_1))
330
+ return output
331
+
332
+ def save_vocabulary(self, save_directory, filename_prefix: Optional[str] = None):
333
+ if not os.path.isdir(save_directory):
334
+ raise ValueError("save directory does not exist")
335
+ filename = (f"{filename_prefix}-" if filename_prefix else "") + "tokenizer.json"
336
+ destination = os.path.join(save_directory, filename)
337
+ if os.path.abspath(destination) != os.path.abspath(self.tokenizer_file):
338
+ shutil.copyfile(self.tokenizer_file, destination)
339
+ return (destination,)
340
+
341
+
342
+ __all__ = ["Rwkv7Tokenizer"]
343
+
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "0": {
4
+ "content": "<|endoftext|>",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ }
11
+ },
12
+ "auto_map": {
13
+ "AutoTokenizer": [
14
+ "tokenization_rwkv7.Rwkv7Tokenizer",
15
+ null
16
+ ]
17
+ },
18
+ "backend": "rwkv_world_trie",
19
+ "eos_token": "<|endoftext|>",
20
+ "pad_token": "<|endoftext|>",
21
+ "padding_side": "left",
22
+ "tokenizer_class": "Rwkv7Tokenizer",
23
+ "unk_token": "<|endoftext|>"
24
+ }