Polish ModernBERT: The Long and Short of Polish Language Understanding
Abstract
Polish ModernBERT introduces efficient Polish encoder variants with extended context that outperform existing Polish BERT-style models across diverse tasks while reducing parameters and inference costs.
Encoder-only Transformers remain effective for discriminative and representation-learning tasks, yet Polish encoders still largely rely on BERT/RoBERTa-style architectures. We introduce Polish ModernBERT, a family of four Polish encoders available at Base and Large scales, each with 512-token and 8K context variants. We adapt the ModernBERT pretraining recipe through staged selection experiments and release a long-context benchmark covering legal topic classification, ideological decision-direction prediction, factual-consistency assessment over literary plot summaries, and human-rights violation assessment. Across 30 tasks, Polish ModernBERT achieves the best overall performance among the evaluated Polish encoders, reaching 83.99 and 85.11 for the Base-8K and Large-8K models, respectively. On long-context tasks, the 8K variants improve over matched Polish RoBERTa-8K baselines from 67.47 to 77.15 and from 75.88 to 78.49 at the Base and Large scales, respectively. The Base-8K model achieves this gain with 22\% fewer parameters (149M vs.\ 190M). Efficiency measurements in representative inference setups show lower peak memory usage and latency than matched Polish RoBERTa baselines in both 512-token and 8K settings. Polish ModernBERT-8K-Base additionally achieves the best result on a Polish retrieval benchmark among the evaluated encoders below 300M parameters.
Models citing this paper 4
OPI-PIB/pl-ModernBERT-base
Datasets citing this paper 5
mmichall/ECtHR-PL-VA
mmichall/SCOTUS-Dom
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper