ChineseBabyLM Treatment NFKC BERT MLM 4.66M

Small from-scratch BERT-style masked language model trained as the Round 20 DeepScientist NFKC-normalization treatment pilot.

Model Details

  • Architecture: BertForMaskedLM
  • Model type: masked language model / encoder-only BERT
  • Parameters: 4,655,803
  • Layers: 4
  • Attention heads: 4
  • Hidden size: 256
  • Vocabulary size: 5,307
  • Max position embeddings: 256

Training Data

Chinese BabyLM official 10k stratified sample with NFKC text normalization.

Evaluation

Official zero-shot evaluator results:

Task Accuracy
ZhoBLiMP 53.53
Hanzi structure 58.00
Hanzi pinyin 74.00
Mean available tasks 61.84

Package Source

This model was packaged from treatment_round20_nfkc_bert_mlm_4p66m in the local ChineseBabyLM best-model archive.

Downloads last month
6
Safetensors
Model size
4.66M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support