PaxiAI commited on
Commit
9dab515
·
verified ·
1 Parent(s): 6444ecf

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -14,9 +14,9 @@ tags:
14
  license: apache-2.0
15
  ---
16
 
17
- # VietToken
18
 
19
- `VietToken` is a **48,000-token Byte-level BPE tokenizer** designed primarily for Vietnamese language models trained from scratch.
20
 
21
  The tokenizer is optimized for Vietnamese while retaining practical coverage of English and source code. It is intended to be architecture-independent and can be used with Qwen-style, LLaMA-style, or other autoregressive language model architectures as long as the model configuration uses the same vocabulary and token IDs.
22
 
@@ -90,7 +90,7 @@ The total serialized corpus size was approximately **8 GiB**.
90
  from transformers import AutoTokenizer
91
 
92
  tokenizer = AutoTokenizer.from_pretrained(
93
- "PaxiAI/VietToken",
94
  use_fast=True,
95
  )
96
 
 
14
  license: apache-2.0
15
  ---
16
 
17
+ # Vietnamese-Tokenizer
18
 
19
+ `Vietnamese-Tokenizer` is a **48,000-token Byte-level BPE tokenizer** designed primarily for Vietnamese language models trained from scratch.
20
 
21
  The tokenizer is optimized for Vietnamese while retaining practical coverage of English and source code. It is intended to be architecture-independent and can be used with Qwen-style, LLaMA-style, or other autoregressive language model architectures as long as the model configuration uses the same vocabulary and token IDs.
22
 
 
90
  from transformers import AutoTokenizer
91
 
92
  tokenizer = AutoTokenizer.from_pretrained(
93
+ "PaxiAI/Vietnamese-Tokenizer",
94
  use_fast=True,
95
  )
96