--- language: - vi - en library_name: transformers pipeline_tag: text-generation tags: - tokenizer - byte-level-bpe - vietnamese - english - code - qwen-style license: apache-2.0 --- # Vietnamese-Tokenizer `Vietnamese-Tokenizer` is a **48,000-token Byte-level BPE tokenizer** designed primarily for Vietnamese language models trained from scratch. The tokenizer is optimized for Vietnamese while retaining practical coverage of English and source code. It is intended to be architecture-independent and can be used with Qwen-style, LLaMA-style, or other autoregressive language model architectures as long as the model configuration uses the same vocabulary and token IDs. ## Key Features - **Vocabulary size:** 48,000 - **Algorithm:** Byte-level BPE - **Unicode normalization:** NFC - **Lowercasing:** No - **Unknown token:** None - **Fast tokenizer:** Yes - **Byte-level coverage:** Any UTF-8 text can be represented - **Vietnamese-focused vocabulary** - Includes English and source-code coverage - Includes reserved token IDs for future extensions ## Special Tokens | Token | ID | |---|---:| | `<|pad|>` | 0 | | `<|bos|>` | 1 | | `<|eos|>` | 2 | | `<|im_start|>` | 3 | | `<|im_end|>` | 4 | | `<|fim_prefix|>` | 5 | | `<|fim_middle|>` | 6 | | `<|fim_suffix|>` | 7 | The tokenizer also reserves IDs `8..263` as: ```text <|reserved_0|> ... <|reserved_255|> ``` These reserved tokens are intentionally kept stable so future model variants can introduce additional control tokens without changing the existing token-to-ID mapping. ## Training Corpus The tokenizer was trained on approximately **8 GiB** of mixed-domain text. Approximate composition: | Domain | Share | |---|---:| | Vietnamese news | 70% | | Vietnamese Wikipedia | 15% | | English general text | 12% | | Source code | 3% | The source-code portion includes multiple programming languages, with higher weight assigned to commonly used languages such as Python, JavaScript, Java, C++, and Go. The corpus used to train this tokenizer is published separately. ## Corpus Statistics The tokenizer training corpus contained approximately: - **1,687,254** Vietnamese news documents - **1,035,370** Vietnamese Wikipedia documents - **326,223** English documents - **~107K** code documents across multiple programming languages The total serialized corpus size was approximately **8 GiB**. ## Usage ```python from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained( "PaxiAI/Vietnamese-Tokenizer", use_fast=True, ) text = "Trí tuệ nhân tạo đang thay đổi cách con người làm việc." ids = tokenizer.encode( text, add_special_tokens=False, ) print(ids) print(tokenizer.decode(ids)) ``` ## Chat Template The tokenizer includes a simple ChatML-style template: ```text <|im_start|>system You are a helpful assistant. <|im_end|> <|im_start|>user Xin chào! <|im_end|> <|im_start|>assistant Chào bạn! <|im_end|> ``` Example: ```python messages = [ { "role": "system", "content": "Bạn là một trợ lý hữu ích." }, { "role": "user", "content": "Xin chào!" } ] prompt = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, ) print(prompt) ``` ## Unicode Handling The tokenizer applies **Unicode NFC normalization**. For example, canonically equivalent NFC and NFD forms of Vietnamese text are normalized to the same representation before tokenization. This is particularly important for Vietnamese because accented characters may otherwise appear in multiple Unicode representations. ## Validation Before release, the tokenizer passed a production validation suite covering: - fixed-text encode/decode round trips - Vietnamese NFC/NFD normalization - special-token ID stability - tokenizer save/reload consistency - chat-template rendering - randomized Unicode fuzz testing - no unknown-token output - long-input tokenization - fast-tokenizer loading - exact vocabulary-size validation A fuzz test of **10,000 randomly generated Unicode strings** completed with zero failures. The tokenizer also produced zero unknown tokens during validation. ## Performance Notes On manual Vietnamese examples, the tokenizer typically produced roughly **3.6-4.4 characters per token**, depending on the sentence. Example: ```text Nguyễn Thị Kim Hương đang được cấp cứu tại Bệnh viện Chợ Rẫy. ``` was tokenized into approximately one token per common Vietnamese syllable or word component. The tokenizer is intentionally optimized more strongly for Vietnamese than for code identifiers or rare English terms. ## Compatibility Warning Once a model has been pretrained with this tokenizer, the following must remain unchanged: - vocabulary - token-to-ID mapping - normalization rules - special-token IDs Changing any of these creates a different tokenizer and should be released under a new version. ## Intended Use This tokenizer is suitable for: - Vietnamese language model pretraining - bilingual Vietnamese-English language models - Vietnamese instruction-tuned models - small and medium autoregressive language models - experimental Qwen-style or LLaMA-style architectures - models with limited parameter budgets where a very large vocabulary would be inefficient ## Limitations - The training corpus is heavily weighted toward Vietnamese news and Wikipedia. - Conversational Vietnamese is less represented than formal written Vietnamese. - Code represents only a small portion of the tokenizer training data. - The tokenizer is not intended to be optimal for multilingual models covering many languages equally. - A tokenizer alone does not determine model quality; pretraining data and training methodology remain critical. ## License This repository contains the tokenizer artifact itself. Please ensure that the selected repository license is compatible with the licensing and redistribution requirements of the tokenizer artifact and its training-data sources.