Model Card for our Cantonese BERT(s)
These models are my final project results in the course IEMS5726A Data Science in Practice during the 2025 Spring term @ CUHK.
Two Models: Bert-ManCan & Bert-ManSikuCan
One Cantonese Corpus: Over 30 Million Cantonese Sentences ๐๐
There are two models in this project, all trained on my own collected Cantonese corpus with over 30 million sentences (see yue_corpus_all_v1.txt).
- Bert-ManCan
This one is continued pre-training based on BERT-base-chinese with our Cantonese corpus data (yue_corpus_all_v1.txt). (bertbasechinese_add_cantonese_mlm_512_checkpoints/final_man_cantonese_bert_512)
This model is called Bert-ManCan.
ManCan is the abbreviation of "Mandarin + Cantonese".
- Bert-ManSikuCan
This one is continued pre-training based on SikuBERT with our Cantonese corpus data (yue_corpus_all_v1.txt). (sikubert_add_cantonese_mlm_512_checkpoints/final_siku_cantonese_bert_512)
This model is called Bert-ManSikuCan.
ManSikuCan is the abbreviation of "Mandarin + SikuQuanshu + Cantonese".
SikuBERT is another model continued pre-training based on BERT-bert-chinese with Siku-Quanshu data. (for more details of this model, see https://huggingface.co/SIKU-BERT/sikubert)
Thus, the only difference between Bert-ManCan and Bert-ManSikuCan on the data level is the additional ancient Chinese SiKuQuanShu corpus in Bert-ManSikuCan, and these two models have exactly the same architecture.
These two models also have differences in vocabulary size, which will be discussed later. The following Table clearly shows the difference of these two models on the data level.
| Model structure | bert-base | bert-base |
|---|---|---|
| Pretraining data / Model name at this stage | Mandarin Wikipedia / bert-base-chinese | Mandarin Wikipedia / bert-base-chinese |
| Continued pretraing data / Model name at this stage | NA / Same model as last stage | SiKuQuanShu / SikuBERT |
| Continued pretraing data / Model name at this stage | Our Cantonese Corpus / Bert-ManCan | Our Cantonese Corpus / Bert-ManSikuCan |
SOTA Performance ๐โ
Our model Bert-ManCan gets SOTA performance (accuracy = 66.48%, previous SOTA accuracy = 60.8%, we have a nearly 6% improvement!๐) on the Cantonese downstream task: Openrice_Cantonese (https://github.com/Christainx/Dataset_Cantonese_Openrice/tree/master) (Xiang et al., 2019).
Openrice_Cantonese is a labelled dataset with 61610 Cantonese comments and corresponding ratings (ranking from 1 to 5) towards restaurants
| Model name | Accuracy | Model name | Accuracy | |
|---|---|---|---|---|
| Models with continued pretraining on Cantonese | Bert-ManCan | 66.48% | Bert-ManSikuCan | 64.78% |
| Models that our continued pretraining based on | bert-base-chinese | 64.99% | SikuBERT | 64.05% |
Before our experiment, the SOTA performance on this dataset is 60.8% by the LSTM-SAT (COMB) algorithm in Xiang et al. (2019).
The improvement from bert-base-chinese of 64.99% to Bert-ManCan of 66.48% is also significant with an improvement of about 1.5%.
Our results clearly show the effectiveness of our Cantonese continued pretraining on downstream Cantonese tasks.
Open-source Checkpoints
To promote the research on training dynamics and language acquisition, we provide the checkpoints between
- the base model BERT-bert-chinese and Bert-ManCan
- bertbasechinese_add_cantonese_mlm_512_checkpoints/checkpoint-[10000/20000/30000/40000/50000/59949]
and
- the base model SikuBERT and Bert-ManSikuCan
- sikubert_add_cantonese_mlm_512_checkpoints/checkpoint-[10000/20000/30000/40000/50000/59949].
Model Details
Model Description
- Developed by: Zijie ZHANG
- Funded by [optional]: [More Information Needed]
- Shared by [optional]: [More Information Needed]
- Model type: BERT
- Language(s) (NLP): Chinese ๆผข่ช/ไธญๆ (include Mandarin ๆฎ้่ฉฑ, Cantonese ๅปฃๅบ็ฒต่ช, Ancient Chinese ๆ่จๆ)
- License: [More Information Needed]
- Continued pre-trained from model [optional]: BERT-base-Chinese & SikuBERT
Training Details
Training Data
yue_corpus_all_v1.txt
We collect 30695134 Cantonese sentences in total from three data sources:
- one is from the corpora mentioned in PyCantonese (a Python library for Cantonese NLP, https://pycantonese.org/data.html) (Lee et al., 2022) (12305 sentences in total after preprocessing),
- one is from Cantonese Wikipedia (zh_yuewiki-latest-pages-articles.xml downloaded from https://dumps.wikimedia.org/zh_yuewiki/ by 2025/04/19) (531842 sentences in total after preprocessing),
- and another one is from https://huggingface.co/datasets/raptorkwok/cantonese_sentences (30150987 sentences in total).
PyCantonese mentioned eight corpora, with one (HKCanCor) built in and seven external corpora, and we use and download six: Child Heritage Chinese Corpus, Guthrie Bilingual Corpus, HKU-70 Corpus, Lee-Wong-Leung Corpus, Leo Corpus, and the Yip-Matthews Bilingual Corpus. We do not use the Paidologos Corpus: Cantonese because it is written in phonetic alphabets rather than the common understanble Chinese characters. The raw data in all external corpora are in CHAT format, which is a common corpus format in Linguistics. Some downloaded external corpora are multilingual or mix-coded, and we exclude all CHAT files in non-Cantonese folders or labelled without Cantonese. We use the .read_chat() function in PyCantonese to read the CHAT files. The built-in HKCanCor corpus is read by the function .hkcancor(). The punctuation marks in the original corpora are all half-angle, which is not consistent with the real Cantonese usage, so we randomly replace 80% of the half-angle punctuation marks with full-angle punctuation marks. Since all corpora mentioned here are oral corpora: Oral Cantonese written in Chinese characters, many utterances are quite short, like "ๅพๅๅพ๏ผ", we concatenate utterances consequently in each CHAT file to sentences with a length of just over 256. We do this concatenation only within each CHAT file because consecutive utterances in one CHAT file are semantically coherent, and utterances between different CHAT files are unrelated.
For the Cantonese Wikipedia, we use the WikiExtractor to extract the pure text from the downloaded xml file. Then, we implement further clean of the pure text. We clean all brackets without any content and all lines without any Wikipedia content. For the lines ending with '.', they are titles or sub-titles on Wikipedia. These lines are usually short and semantically closely related to the next line, which is often the detailed description of the title/sub-title, thus we concatenate each title/sub-title line with its following line and replacing the original '.' with '๏ผ', which is the common punctuation mark between titles/sub-titles and their description details in daily written text, to make a sentence. We also find that some lines are begin with '๏ผ', which are consecutive to the last line above, so we concatenate each line beginning with '๏ผ' with its last line above.
The data downloaded from https://huggingface.co/datasets/raptorkwok/cantonese_sentences are all in Parquet format, which is a data format very suitable for big data. We use pandas.read_parquet() to read these files and extract the pure text.
We set the maximum sequence length in our BERT to 512, and for all the text lines got above with a length longer than 512, we split them into several consecutive sentences with a length of 512.
Through the data collection and preprocessing procedures, we get 30695134 sentences in total and store them in a txt file named yue_corpus_all_v1.txt line by line.
[More Information Needed]
Training Procedure
Preprocessing
For the two base models, bert-base-chinese and SikuBERT, we download them from the Transformer library with codes AutoTokenizer.from_pretrained() and AutoModelForMaskedLM.from_pretrained().
Then, we extend the vocabulary of these two models because some Chinese characters used in Cantonese are not used in Mandarin. If not doing so, these Chinese characters will be tokened as [UNK] in the model. For all the characters in yue_corpus_all_v1.txt but not in corresponding model vocabulary (the SikuBERT vocabulary is already an extended one based on bert-base-chinese), we add them (denoted by new_token) to the vocabulary by BertTokenizer.from_pretrained(source_model_path).add_tokens, along with the '##'+new_token version also added, because in bert-base-chinese vocabulary all character tokens also have a version with prefix '##'. As a result, the Bert-ManCan has a vocabulary size of 51430 and the Bert-ManSikuCan has a vocabulary size of 47735.
After extending the vocabularies of these two models, these models are updated by model.resize_token_embeddings(new_vocabulary_size) because BERT's token embedding layer has a size of [vocab_size, hidden_size]. The dimension of the token embedding layer should be consistent with the vocabulary size.
Training Hyperparameters
The BERT model (Delvin et al., 2018) is pretrained on a large-scale unlabelled text corpus in the manner of Masked Language Model (MLM) and Next Sentence Prediction (NSP). Our continued pretraining is only in the manner of MLM due to the following reasons:
โ Liu et al. (2019) proved that NSP does not benefit the model performance on downstream tasks,
โก NSP need clear boundaries between different documents to know which sentences are semantically coherent, but more than 98% of our data, which are from HuggingFace, do not have document boundaries,
โข one of our base model โโ SikuBERT โโ is also continued pretrained from bert-base-chinese without NSP and we want to keep this consistency.
We use two A40 GPUs (provided by CUHK Department of Information Engineering, thanks a lot๐โ!), each with nearly 48GB vRAM, and due to the time schedule and deadline of the project, we set a limitation of training these two models in total in two days and using about half vRAM resource for each GPU (because the GPU server is shared by a research group and we do not want to monopoly).
We finally achieve this goal by training each model in each single GPU simultaneously, each model training occupies about 23GB vRAM, and Bert-ManCan training costs 30 hours and 54 minutes, and Bert-ManSikuCan training costs 32 hours and 28 minutes with the following hyper-parameters (exactly the same for the training of these two models):
training_args = TrainingArguments(
output_dir=OUTPUT_DIR,
overwrite_output_dir=True,
num_train_epochs=1,
per_device_train_batch_size=64,
gradient_accumulation_steps=8,
warmup_ratio=0.1,
learning_rate=6e-5,
weight_decay=0.003,
fp16=True,
fp16_opt_level="O2",
fp16_full_eval=True,
optim="adamw_torch_fused",
max_grad_norm=1.0,
no_cuda=False,
gradient_checkpointing=True,
dataloader_num_workers=32,
dataloader_pin_memory=True,
dataloader_prefetch_factor=4,
remove_unused_columns=True,
save_steps=10000,
logging_steps=2000,
report_to="tensorboard",
prediction_loss_only=True,
log_level="info",
)
We set epoch to 1 because
โ we implement the continued pretraining rather than training from scratch, which means that our base models already have the basic language ability,
โก we don't want our training destroy the models' language abilities on Mandarin (and ancient Chinese for Bert-ManSikuCan), and
โข larger epoch means longer training time, or larger vRAM usage with larger batch size, which can exceed our expectation.
Here shows the training loss change of these two models:
