DialLM CPT π
Collection
Continual pre-training checkpoints using ICE for DialLM across Gemma, Llama, and Qwen base models. β’ 3 items β’ Updated
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main).
Continually pretrained on the International Corpus of English (18 varieties, ~20M tokens) using GaLore for memory-efficient full-parameter optimisation. This is the shared foundation checkpoint both the implicit and explicit adaptation threads branch from.
Code, checkpoints, preference datasets, linguistic-analysis toolkit: https://github.com/surrey-nlp/diallm
Paper: https://arxiv.org/abs/2607.07669
@article{painter2026diallm,
title = {DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation},
author = {Painter, Jordan and Srirag, Dipankar and Kappiyath, Adarsh and Kanojia, Diptesh and Joshi, Aditya and Yin, Lu},
year = {2026},
eprint = {2607.07669},
archivePrefix = {arXiv}
}