DiaLLM β€” Qwen 3-8B β€” CPT

Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main).

DiaLLM pipeline

  • Base model: Qwen 3-8B
  • Pipeline stage: continual pretraining (CPT), precedes dialect-specific branching

Continually pretrained on the International Corpus of English (18 varieties, ~20M tokens) using GaLore for memory-efficient full-parameter optimisation. This is the shared foundation checkpoint both the implicit and explicit adaptation threads branch from.

Code, checkpoints, preference datasets, linguistic-analysis toolkit: https://github.com/surrey-nlp/diallm

Paper: https://arxiv.org/abs/2607.07669

Citation

@article{painter2026diallm,
  title     = {DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation},
  author    = {Painter, Jordan and Srirag, Dipankar and Kappiyath, Adarsh and Kanojia, Diptesh and Joshi, Aditya and Yin, Lu},
  year      = {2026},
  eprint    = {2607.07669},
  archivePrefix = {arXiv}
}
Downloads last month
20
Safetensors
Model size
8B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for jordanpainter/diallm-qwen-cpt

Finetuned
Qwen/Qwen3-8B
Finetuned
(2020)
this model
Finetunes
1 model

Collection including jordanpainter/diallm-qwen-cpt

Papers for jordanpainter/diallm-qwen-cpt