Model Card for our Cantonese BERT(s)

These models are my final project results in the course IEMS5726A Data Science in Practice during the 2025 Spring term @ CUHK.

Two Models: Bert-ManCan & Bert-ManSikuCan

One Cantonese Corpus: Over 30 Million Cantonese Sentences ๐Ÿ“–๐Ÿš€

There are two models in this project, all trained on my own collected Cantonese corpus with over 30 million sentences (see yue_corpus_all_v1.txt).

  • Bert-ManCan

This one is continued pre-training based on BERT-base-chinese with our Cantonese corpus data (yue_corpus_all_v1.txt). (bertbasechinese_add_cantonese_mlm_512_checkpoints/final_man_cantonese_bert_512)

This model is called Bert-ManCan.

ManCan is the abbreviation of "Mandarin + Cantonese".

  • Bert-ManSikuCan

This one is continued pre-training based on SikuBERT with our Cantonese corpus data (yue_corpus_all_v1.txt). (sikubert_add_cantonese_mlm_512_checkpoints/final_siku_cantonese_bert_512)

This model is called Bert-ManSikuCan.

ManSikuCan is the abbreviation of "Mandarin + SikuQuanshu + Cantonese".

SikuBERT is another model continued pre-training based on BERT-bert-chinese with Siku-Quanshu data. (for more details of this model, see https://huggingface.co/SIKU-BERT/sikubert)

Thus, the only difference between Bert-ManCan and Bert-ManSikuCan on the data level is the additional ancient Chinese SiKuQuanShu corpus in Bert-ManSikuCan, and these two models have exactly the same architecture.

These two models also have differences in vocabulary size, which will be discussed later. The following Table clearly shows the difference of these two models on the data level.

Model structure bert-base bert-base
Pretraining data / Model name at this stage Mandarin Wikipedia / bert-base-chinese Mandarin Wikipedia / bert-base-chinese
Continued pretraing data / Model name at this stage NA / Same model as last stage SiKuQuanShu / SikuBERT
Continued pretraing data / Model name at this stage Our Cantonese Corpus / Bert-ManCan Our Cantonese Corpus / Bert-ManSikuCan

SOTA Performance ๐Ÿš€โœŒ

Our model Bert-ManCan gets SOTA performance (accuracy = 66.48%, previous SOTA accuracy = 60.8%, we have a nearly 6% improvement!๐Ÿš€) on the Cantonese downstream task: Openrice_Cantonese (https://github.com/Christainx/Dataset_Cantonese_Openrice/tree/master) (Xiang et al., 2019).

Openrice_Cantonese is a labelled dataset with 61610 Cantonese comments and corresponding ratings (ranking from 1 to 5) towards restaurants

Model name Accuracy Model name Accuracy
Models with continued pretraining on Cantonese Bert-ManCan 66.48% Bert-ManSikuCan 64.78%
Models that our continued pretraining based on bert-base-chinese 64.99% SikuBERT 64.05%

Before our experiment, the SOTA performance on this dataset is 60.8% by the LSTM-SAT (COMB) algorithm in Xiang et al. (2019).

The improvement from bert-base-chinese of 64.99% to Bert-ManCan of 66.48% is also significant with an improvement of about 1.5%.

Our results clearly show the effectiveness of our Cantonese continued pretraining on downstream Cantonese tasks.

Open-source Checkpoints

To promote the research on training dynamics and language acquisition, we provide the checkpoints between

  • the base model BERT-bert-chinese and Bert-ManCan
  • bertbasechinese_add_cantonese_mlm_512_checkpoints/checkpoint-[10000/20000/30000/40000/50000/59949]

and

  • the base model SikuBERT and Bert-ManSikuCan
  • sikubert_add_cantonese_mlm_512_checkpoints/checkpoint-[10000/20000/30000/40000/50000/59949].

Model Details

Model Description

  • Developed by: Zijie ZHANG
  • Funded by [optional]: [More Information Needed]
  • Shared by [optional]: [More Information Needed]
  • Model type: BERT
  • Language(s) (NLP): Chinese ๆผข่ชž/ไธญๆ–‡ (include Mandarin ๆ™ฎ้€š่ฉฑ, Cantonese ๅปฃๅบœ็ฒต่ชž, Ancient Chinese ๆ–‡่จ€ๆ–‡)
  • License: [More Information Needed]
  • Continued pre-trained from model [optional]: BERT-base-Chinese & SikuBERT

Training Details

Training Data

yue_corpus_all_v1.txt

We collect 30695134 Cantonese sentences in total from three data sources:

PyCantonese mentioned eight corpora, with one (HKCanCor) built in and seven external corpora, and we use and download six: Child Heritage Chinese Corpus, Guthrie Bilingual Corpus, HKU-70 Corpus, Lee-Wong-Leung Corpus, Leo Corpus, and the Yip-Matthews Bilingual Corpus. We do not use the Paidologos Corpus: Cantonese because it is written in phonetic alphabets rather than the common understanble Chinese characters. The raw data in all external corpora are in CHAT format, which is a common corpus format in Linguistics. Some downloaded external corpora are multilingual or mix-coded, and we exclude all CHAT files in non-Cantonese folders or labelled without Cantonese. We use the .read_chat() function in PyCantonese to read the CHAT files. The built-in HKCanCor corpus is read by the function .hkcancor(). The punctuation marks in the original corpora are all half-angle, which is not consistent with the real Cantonese usage, so we randomly replace 80% of the half-angle punctuation marks with full-angle punctuation marks. Since all corpora mentioned here are oral corpora: Oral Cantonese written in Chinese characters, many utterances are quite short, like "ๅพ—ๅ””ๅพ—๏ผŸ", we concatenate utterances consequently in each CHAT file to sentences with a length of just over 256. We do this concatenation only within each CHAT file because consecutive utterances in one CHAT file are semantically coherent, and utterances between different CHAT files are unrelated.

For the Cantonese Wikipedia, we use the WikiExtractor to extract the pure text from the downloaded xml file. Then, we implement further clean of the pure text. We clean all brackets without any content and all lines without any Wikipedia content. For the lines ending with '.', they are titles or sub-titles on Wikipedia. These lines are usually short and semantically closely related to the next line, which is often the detailed description of the title/sub-title, thus we concatenate each title/sub-title line with its following line and replacing the original '.' with '๏ผš', which is the common punctuation mark between titles/sub-titles and their description details in daily written text, to make a sentence. We also find that some lines are begin with '๏ผŒ', which are consecutive to the last line above, so we concatenate each line beginning with '๏ผŒ' with its last line above.

The data downloaded from https://huggingface.co/datasets/raptorkwok/cantonese_sentences are all in Parquet format, which is a data format very suitable for big data. We use pandas.read_parquet() to read these files and extract the pure text.

We set the maximum sequence length in our BERT to 512, and for all the text lines got above with a length longer than 512, we split them into several consecutive sentences with a length of 512.

Through the data collection and preprocessing procedures, we get 30695134 sentences in total and store them in a txt file named yue_corpus_all_v1.txt line by line.

[More Information Needed]

Training Procedure

Preprocessing

For the two base models, bert-base-chinese and SikuBERT, we download them from the Transformer library with codes AutoTokenizer.from_pretrained() and AutoModelForMaskedLM.from_pretrained().

Then, we extend the vocabulary of these two models because some Chinese characters used in Cantonese are not used in Mandarin. If not doing so, these Chinese characters will be tokened as [UNK] in the model. For all the characters in yue_corpus_all_v1.txt but not in corresponding model vocabulary (the SikuBERT vocabulary is already an extended one based on bert-base-chinese), we add them (denoted by new_token) to the vocabulary by BertTokenizer.from_pretrained(source_model_path).add_tokens, along with the '##'+new_token version also added, because in bert-base-chinese vocabulary all character tokens also have a version with prefix '##'. As a result, the Bert-ManCan has a vocabulary size of 51430 and the Bert-ManSikuCan has a vocabulary size of 47735.

After extending the vocabularies of these two models, these models are updated by model.resize_token_embeddings(new_vocabulary_size) because BERT's token embedding layer has a size of [vocab_size, hidden_size]. The dimension of the token embedding layer should be consistent with the vocabulary size.

Training Hyperparameters

The BERT model (Delvin et al., 2018) is pretrained on a large-scale unlabelled text corpus in the manner of Masked Language Model (MLM) and Next Sentence Prediction (NSP). Our continued pretraining is only in the manner of MLM due to the following reasons:

โ‘  Liu et al. (2019) proved that NSP does not benefit the model performance on downstream tasks,

โ‘ก NSP need clear boundaries between different documents to know which sentences are semantically coherent, but more than 98% of our data, which are from HuggingFace, do not have document boundaries,

โ‘ข one of our base model โ€”โ€” SikuBERT โ€”โ€” is also continued pretrained from bert-base-chinese without NSP and we want to keep this consistency.

We use two A40 GPUs (provided by CUHK Department of Information Engineering, thanks a lot๐Ÿ™‡โ€!), each with nearly 48GB vRAM, and due to the time schedule and deadline of the project, we set a limitation of training these two models in total in two days and using about half vRAM resource for each GPU (because the GPU server is shared by a research group and we do not want to monopoly).

We finally achieve this goal by training each model in each single GPU simultaneously, each model training occupies about 23GB vRAM, and Bert-ManCan training costs 30 hours and 54 minutes, and Bert-ManSikuCan training costs 32 hours and 28 minutes with the following hyper-parameters (exactly the same for the training of these two models):

training_args = TrainingArguments(

output_dir=OUTPUT_DIR,

overwrite_output_dir=True,

num_train_epochs=1,

per_device_train_batch_size=64,

gradient_accumulation_steps=8,

warmup_ratio=0.1,

learning_rate=6e-5,

weight_decay=0.003,

fp16=True,  

fp16_opt_level="O2",

fp16_full_eval=True,  

optim="adamw_torch_fused",  

max_grad_norm=1.0,

no_cuda=False, 

gradient_checkpointing=True,

dataloader_num_workers=32,

dataloader_pin_memory=True,

dataloader_prefetch_factor=4,

remove_unused_columns=True,

save_steps=10000,

logging_steps=2000,

report_to="tensorboard",

prediction_loss_only=True,

log_level="info",

)

We set epoch to 1 because

โ‘  we implement the continued pretraining rather than training from scratch, which means that our base models already have the basic language ability,

โ‘ก we don't want our training destroy the models' language abilities on Mandarin (and ancient Chinese for Bert-ManSikuCan), and

โ‘ข larger epoch means longer training time, or larger vRAM usage with larger batch size, which can exceed our expectation.

Here shows the training loss change of these two models:

CantoneseBERT_loss_curve

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Datasets used to train zijie-zhang/CantoneseBERT