Instructions to use OPI-PIB/pl-ModernBERT-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OPI-PIB/pl-ModernBERT-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="OPI-PIB/pl-ModernBERT-large")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("OPI-PIB/pl-ModernBERT-large") model = AutoModelForMaskedLM.from_pretrained("OPI-PIB/pl-ModernBERT-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Polish ModernBERT Large
Polish ModernBERT is a family of monolingual ModernBERT-based encoders pretrained for Polish. The family covers two model scales (Base and Large) and two context lengths (512 and 8,192 tokens). The models are general-purpose pretrained encoders intended to serve as strong foundations for a broad range of Polish NLP applications. They can be fine-tuned or adapted for downstream tasks including, but not limited to, text and document classification, sequence labeling, regression, semantic similarity, retrieval, reranking, and representation learning.
📄 Paper: Polish ModernBERT: The Long and Short of Polish Language Understanding
🤗 Model collection: Polish ModernBERT
📚 LongContext benchmark: Dataset
Model family
| Model | Parameters | Context | Vocabulary | Tokenizer |
|---|---|---|---|---|
pl-ModernBERT-512-base |
149M | 512 | 50,008 | SentencePiece Unigram |
pl-ModernBERT-512-large |
475M | 512 | 128,256 | SentencePiece Unigram + byte fallback |
pl-ModernBERT-base |
149M | 8,192 | 50,008 | SentencePiece Unigram |
pl-ModernBERT-large |
475M | 8,192 | 128,256 | SentencePiece Unigram + byte fallback |
All variants use the core ModernBERT architecture with RoPE, GeGLU feed-forward layers, pre-normalization, and alternating global/local attention. The Base models use 22 layers with hidden size 768, while the Large models use 28 layers with hidden size 1,024.
Training
The models were pretrained on approximately 44.5B Polish tokens (197 GB after preprocessing) from a curated Polish corpus, Common Crawl (CC-MAIN-2019-43), and the Polish pol_Latn subset of FineTranslations.
Pretraining follows a staged recipe selected through downstream validation: four 512-token stages progressively transition from token-level MLM to whole-word masking, increased emphasis on curated data, and final annealing, followed by long-context continuation to 8,192 tokens. During context extension, the global RoPE theta is increased from 10,000 to 160,000.
Training was performed with Composer in BF16 using StableAdamW and distributed data parallelism across 8 NVIDIA GH200 GPUs.
Evaluation
The models were evaluated on 30 Polish NLU tasks spanning KLEJ, FinBench, the five-task LongContext benchmark, and a range of additional Polish NLP tasks. Each model-task configuration was fine-tuned using five random seeds; scores below are means on a 0–100 scale.
Evaluation Results
The tables below summarize average performance across the main evaluation groups. Scores are reported on a 0–100 scale. Overall Avg. is the macro-average across all 30 individual tasks rather than the average of the four group-level scores.
Base models
| Model | Context | KLEJ Avg. | FinBench Avg. | Other Tasks Avg. | LongContext Avg. | Overall Avg. |
|---|---|---|---|---|---|---|
| XLM-R Base | 512 | 84.46 | 83.13 | 77.10 | 50.75 | 76.12 |
| HerBERT Base | 512 | 85.83 | 83.65 | 79.28 | 61.93 | 79.23 |
| Polish RoBERTa-v2 Base | 512 | 86.75 | 85.14 | 80.26 | 57.66 | 79.42 |
| Polish ModernBERT 512 Base | 512 | 86.84 | 86.95 | 82.07 | 71.59 | 82.73 |
| EuroBERT-210M | 8K | 77.16 | 82.38 | 73.71 | 72.30 | 76.24 |
| mmBERT Small | 8K | 80.11 | 83.07 | 78.01 | 69.02 | 78.15 |
| Polish RoBERTa-8K Base | 8K | 86.86 | 85.98 | 82.36 | 67.47 | 81.95 |
| Polish ModernBERT Base | 8K | 87.00 | 86.69 | 83.08 | 77.15 | 83.99 |
Large models
| Model | Context | KLEJ Avg. | FinBench Avg. | Other Tasks Avg. | LongContext Avg. | Overall Avg. |
|---|---|---|---|---|---|---|
| XLM-R Large | 512 | 87.29 | 85.44 | 82.64 | 67.87 | 82.13 |
| HerBERT Large | 512 | 87.83 | 87.33 | 83.42 | 70.26 | 83.33 |
| Polish RoBERTa-v2 Large | 512 | 88.69 | 87.90 | 83.58 | 70.49 | 83.80 |
| Polish ModernBERT 512 Large | 512 | 88.48 | 88.18 | 83.60 | 73.58 | 84.31 |
| EuroBERT-610M | 8K | 80.10 | 86.11 | 78.95 | 77.36 | 80.46 |
| mmBERT Base | 8K | 83.17 | 85.69 | 80.80 | 73.44 | 81.27 |
| Polish RoBERTa-8K Large | 8K | 88.52 | 87.98 | 84.27 | 75.88 | 84.89 |
| Polish ModernBERT Large | 8K | 88.31 | 87.83 | 83.90 | 78.49 | 85.11 |
KLEJ
Polish ModernBERT is competitive with the strongest Polish BERT/RoBERTa encoders on KLEJ. The Base variants obtain the highest KLEJ average in both context settings, while the Large variants remain within 0.21 points of the corresponding Polish RoBERTa models.
Detailed KLEJ results - Base models
| Task | pl-RoBERTa-v2-base | pl-ModernBERT-512-base | pl-RoBERTa-8K-base | pl-ModernBERT-base |
|---|---|---|---|---|
| NKJP-NER | 94.32 | 94.38 | 94.16 | 94.49 |
| CDSC-E | 94.05 | 94.66 | 94.54 | 94.46 |
| CDSC-R | 94.64 | 94.11 | 94.90 | 94.06 |
| CBD | 70.57 | 68.56 | 69.35 | 71.40 |
| POLEMO-IN | 90.97 | 92.88 | 91.27 | 92.14 |
| POLEMO-OUT | 79.11 | 83.77 | 81.26 | 83.04 |
| DYK | 70.38 | 66.90 | 69.35 | 67.28 |
| PSC | 98.88 | 97.79 | 98.90 | 97.68 |
| AR | 87.83 | 88.53 | 88.05 | 88.46 |
| Average | 86.75 | 86.84 | 86.86 | 87.00 |
Detailed KLEJ results - Large models
| Task | pl-RoBERTa-v2-large | pl-ModernBERT-512-large | pl-RoBERTa-8K-large | pl-ModernBERT-large |
|---|---|---|---|---|
| NKJP-NER | 95.75 | 95.05 | 95.64 | 94.38 |
| CDSC-E | 94.16 | 94.60 | 94.28 | 94.68 |
| CDSC-R | 95.25 | 95.14 | 95.33 | 94.47 |
| CBD | 73.10 | 72.60 | 73.23 | 71.47 |
| POLEMO-IN | 93.55 | 93.38 | 93.05 | 93.05 |
| POLEMO-OUT | 83.81 | 84.41 | 83.64 | 84.78 |
| DYK | 74.87 | 73.28 | 74.05 | 74.63 |
| PSC | 98.37 | 98.81 | 98.56 | 98.47 |
| AR | 89.36 | 89.07 | 88.91 | 88.88 |
| Average | 88.69 | 88.48 | 88.52 | 88.31 |
FinBench
Polish ModernBERT shows particularly strong performance on FinBench. The Base variants achieve the highest FinBench average in both context settings. At Large scale, the 512-token Polish ModernBERT obtains the highest average, while the 8K variant remains close to the corresponding Polish RoBERTa model.
Detailed FinBench results - Base models
| Task | pl-RoBERTa-v2-base | pl-ModernBERT-512-base | pl-RoBERTa-8K-base | pl-ModernBERT-base |
|---|---|---|---|---|
| Banking-Short | 78.75 | 80.41 | 79.79 | 80.08 |
| Banking-Long | 85.03 | 87.29 | 86.99 | 87.16 |
| Banking77 | 88.26 | 91.85 | 89.27 | 91.66 |
| FPB | 83.55 | 83.20 | 83.63 | 83.40 |
| GCN | 95.02 | 94.87 | 94.87 | 94.83 |
| Stooq | 80.25 | 84.08 | 81.32 | 83.03 |
| Average | 85.14 | 86.95 | 85.98 | 86.69 |
Detailed FinBench results - Large models
| Task | pl-RoBERTa-v2-large | pl-ModernBERT-512-large | pl-RoBERTa-8K-large | pl-ModernBERT-large |
|---|---|---|---|---|
| Banking-Short | 81.69 | 82.07 | 81.99 | 81.94 |
| Banking-Long | 87.89 | 88.40 | 88.35 | 88.89 |
| Banking77 | 92.45 | 92.96 | 92.74 | 92.62 |
| FPB | 85.26 | 84.80 | 85.42 | 84.60 |
| GCN | 95.04 | 95.08 | 94.97 | 94.88 |
| Stooq | 85.07 | 85.77 | 84.41 | 84.02 |
| Average | 87.90 | 88.18 | 87.98 | 87.83 |
Other Tasks
Across the additional Polish NLP tasks, the Base variants of Polish ModernBERT obtain the highest average in both context settings. At Large scale, the 512-token model achieves the highest average, while the 8K variant remains competitive with the corresponding Polish RoBERTa model.
Detailed Other Tasks results - Base models
| Task | pl-RoBERTa-v2-base | pl-ModernBERT-512-base | pl-RoBERTa-8K-base | pl-ModernBERT-base |
|---|---|---|---|---|
| 8TAGS | 78.03 | 80.69 | 79.21 | 80.86 |
| BAN-PL | 92.19 | 93.10 | 92.62 | 93.08 |
| MIPD | 58.58 | 67.11 | 64.39 | 68.03 |
| PPC | 87.05 | 84.40 | 86.02 | 85.70 |
| SICK-E | 86.71 | 86.31 | 86.31 | 86.61 |
| SICK-R | 82.58 | 83.16 | 83.16 | 83.77 |
| TwitterEMO | 66.46 | 69.02 | 68.75 | 69.52 |
| IMDB | 91.06 | 92.05 | 95.02 | 94.40 |
| EURLEX | 74.51 | 79.31 | 79.12 | 79.61 |
| NKJP-NER* | 85.41 | 85.54 | 88.97 | 89.21 |
| Average | 80.26 | 82.07 | 82.36 | 83.08 |
Detailed Other Tasks results - Large models
| Task | pl-RoBERTa-v2-large | pl-ModernBERT-512-large | pl-RoBERTa-8K-large | pl-ModernBERT-large |
|---|---|---|---|---|
| 8TAGS | 81.64 | 82.50 | 81.44 | 82.24 |
| BAN-PL | 93.80 | 94.00 | 93.99 | 93.51 |
| MIPD | 67.27 | 68.28 | 68.50 | 68.99 |
| PPC | 89.96 | 88.04 | 89.48 | 87.20 |
| SICK-E | 88.33 | 87.88 | 88.96 | 87.47 |
| SICK-R | 85.93 | 84.69 | 86.54 | 84.91 |
| TwitterEMO | 70.70 | 70.20 | 70.60 | 70.35 |
| IMDB | 94.36 | 93.77 | 96.03 | 95.93 |
| EURLEX | 79.19 | 79.84 | 79.77 | 79.76 |
| NKJP-NER* | 84.62 | 86.84 | 87.36 | 88.66 |
| Average | 83.58 | 83.60 | 84.27 | 83.90 |
Long-context performance
Polish ModernBERT shows its largest gains on the LongContext benchmark, which consists of five tasks designed specifically to evaluate long-document understanding. The 8K variants achieve the highest LongContext average at both model scales.
At Base scale, pl-ModernBERT-base improves over pl-RoBERTa-8K-base by 9.68 points (77.15 vs. 67.47) while using 22% fewer parameters (149M vs. 190M). At Large scale, pl-ModernBERT-large improves over the corresponding Polish RoBERTa-8K baseline by 2.61 points (78.49 vs. 75.88).
Detailed LongContext results - Base models
| Task | pl-RoBERTa-v2-base | pl-ModernBERT-512-base | pl-RoBERTa-8K-base | pl-ModernBERT-base |
|---|---|---|---|---|
| SCOTUS-Dom | 79.12 | 82.81 | 79.26 | 84.48 |
| SCOTUS-Dec | 69.86 | 70.74 | 63.20 | 77.79 |
| BookSummary | 85.02 | 83.71 | 88.96 | 90.22 |
| ECtHR-PL-AVA | 33.65 | 61.42 | 64.66 | 68.01 |
| ECtHR-PL-VA | 20.65 | 59.29 | 41.28 | 65.27 |
| Average | 57.66 | 71.59 | 67.47 | 77.15 |
Detailed LongContext results - Large models
| Task | pl-RoBERTa-v2-large | pl-ModernBERT-512-large | pl-RoBERTa-8K-large | pl-ModernBERT-large |
|---|---|---|---|---|
| SCOTUS-Dom | 83.01 | 83.83 | 83.99 | 85.78 |
| SCOTUS-Dec | 72.46 | 71.39 | 68.73 | 78.21 |
| BookSummary | 87.47 | 86.54 | 93.11 | 91.74 |
| ECtHR-PL-AVA | 58.33 | 64.62 | 68.29 | 69.48 |
| ECtHR-PL-VA | 51.17 | 61.50 | 65.27 | 67.24 |
| Average | 70.49 | 73.58 | 75.88 | 78.49 |
The 512-token variants are evaluated on LongContext with truncation and should be treated as practical short-context baselines rather than controlled context-length ablations.
Inference efficiency
In a common BF16 inference setup on a single NVIDIA H100, Polish ModernBERT provides favorable quality–efficiency trade-offs relative to the corresponding Polish RoBERTa baselines.
For 512-token inference, the Base and Large models reduce latency by approximately 26% and 47%, respectively, while reducing peak GPU memory usage by 54% and 18%. In the 8K setting, the corresponding latency reductions are 6% and 25%, with peak memory reductions of 24% and 21%.
Usage
from transformers import AutoTokenizer, AutoModel
model_id = "OPI-PIB/pl-ModernBERT-large"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
text = "W białodrzewiu jaśnie dźni słoneczno, miodzie złoci białopałem żyśnie, drzewia pełni pszczelą i pasieczną, a przez liście kraśnie pęk słowiśnie."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=8192,
)
outputs = model(**inputs)
token_embeddings = outputs.last_hidden_state
For downstream tasks such as classification, regression, sequence labeling, retrieval, or reranking, the pretrained encoder can be adapted to the target task.
The raw checkpoints are not retrieval-specific sentence-embedding models. PIRB retrieval results reported in the paper were obtained after contrastive fine-tuning.
For longer inputs, use one of the corresponding 8K checkpoints from the Polish ModernBERT family.
Flash Attention 2
🏎️ For the highest training and inference efficiency, we recommend using Polish ModernBERT with Flash Attention 2 when supported by your GPU and environment.
pip install flash-attn --no-build-isolation
The model can then be loaded with:
from transformers import AutoModel
model = AutoModel.from_pretrained(
model_id,
attn_implementation="flash_attention_2",
torch_dtype="auto",
)
Authors
Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata, Małgorzata Grębowiec
AI Lab, National Information Processing Institute
(Ośrodek Przetwarzania Informacji – Państwowy Instytut Badawczy, OPI PIB)
Warsaw, Poland
Corresponding author: mperelkiewicz@opi.org.pl
Acknowledgments
This work was supported by the Gaia AI Factory project, funded by the European Union under Grant Agreement No. 101314359 through the EuroHPC Joint Undertaking (EuroHPC JU).
We gratefully acknowledge the Polish high-performance computing infrastructure PLGrid (HPC Center: ACK Cyfronet AGH) for providing computational resources and support within computational grant PLG/2025/018315.
Citation
If you use Polish ModernBERT in your work, please cite:
@misc{perełkiewicz2026polishmodernbertlongshort,
title={Polish ModernBERT: The Long and Short of Polish Language Understanding},
author={Michał Perełkiewicz and Sławomir Dadas and Rafał Poświata and Małgorzata Grębowiec},
year={2026},
eprint={2609.01379},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.01379},
}
- Downloads last month
- 44