PaxiAI commited on
Commit
6444ecf
·
verified ·
1 Parent(s): 05813f4

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +0 -54
README.md CHANGED
@@ -105,16 +105,6 @@ print(ids)
105
  print(tokenizer.decode(ids))
106
  ```
107
 
108
- Example behavior:
109
-
110
- ```text
111
- Input:
112
- Trí tuệ nhân tạo đang thay đổi cách con người làm việc và học tập.
113
-
114
- Decoded:
115
- Trí tuệ nhân tạo đang thay đổi cách con người làm việc và học tập.
116
- ```
117
-
118
  ## Chat Template
119
 
120
  The tokenizer includes a simple ChatML-style template:
@@ -195,36 +185,6 @@ was tokenized into approximately one token per common Vietnamese syllable or wor
195
 
196
  The tokenizer is intentionally optimized more strongly for Vietnamese than for code identifiers or rare English terms.
197
 
198
- ## Architecture Independence
199
-
200
- This tokenizer is not tied to a specific model implementation.
201
-
202
- It can be used with a model initialized from scratch, for example:
203
-
204
- ```python
205
- from transformers import AutoTokenizer, Qwen2Config, Qwen2ForCausalLM
206
-
207
- tokenizer = AutoTokenizer.from_pretrained(
208
- "PaxiAI/VietToken"
209
- )
210
-
211
- config = Qwen2Config(
212
- vocab_size=len(tokenizer),
213
- hidden_size=1024,
214
- intermediate_size=2816,
215
- num_hidden_layers=20,
216
- num_attention_heads=16,
217
- num_key_value_heads=4,
218
- pad_token_id=tokenizer.pad_token_id,
219
- bos_token_id=tokenizer.bos_token_id,
220
- eos_token_id=tokenizer.eos_token_id,
221
- )
222
-
223
- model = Qwen2ForCausalLM(config)
224
- ```
225
-
226
- The model above is randomly initialized. No Qwen weights are required.
227
-
228
  ## Compatibility Warning
229
 
230
  Once a model has been pretrained with this tokenizer, the following must remain unchanged:
@@ -236,20 +196,6 @@ Once a model has been pretrained with this tokenizer, the following must remain
236
 
237
  Changing any of these creates a different tokenizer and should be released under a new version.
238
 
239
- ## Recommended Versioning
240
-
241
- This release should be treated as:
242
-
243
- ```text
244
- VietnameseTokenizer-v1
245
- ```
246
-
247
- Future incompatible tokenizer changes should use a new version such as:
248
-
249
- ```text
250
- VietnameseTokenizer-v2
251
- ```
252
-
253
  ## Intended Use
254
 
255
  This tokenizer is suitable for:
 
105
  print(tokenizer.decode(ids))
106
  ```
107
 
 
 
 
 
 
 
 
 
 
 
108
  ## Chat Template
109
 
110
  The tokenizer includes a simple ChatML-style template:
 
185
 
186
  The tokenizer is intentionally optimized more strongly for Vietnamese than for code identifiers or rare English terms.
187
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
188
  ## Compatibility Warning
189
 
190
  Once a model has been pretrained with this tokenizer, the following must remain unchanged:
 
196
 
197
  Changing any of these creates a different tokenizer and should be released under a new version.
198
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
199
  ## Intended Use
200
 
201
  This tokenizer is suitable for: