Long-Context Encoder Models for Polish Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918384487104512 |
|---|---|
| author | Dadas, Sławomir Poświata, Rafał Kozłowski, Marek Grębowiec, Małgorzata Perełkiewicz, Michał Klimiuk, Paweł Boruta, Przemysław |
| author_facet | Dadas, Sławomir Poświata, Rafał Kozłowski, Marek Grębowiec, Małgorzata Perełkiewicz, Michał Klimiuk, Paweł Boruta, Przemysław |
| contents | While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for discriminative tasks. However, classic encoders like BERT are limited by a short context window, which is insufficient for processing long documents. In this paper, we address this limitation for the Polish by introducing a high-quality Polish model capable of processing sequences of up to 8192 tokens. The model was developed by employing a two-stage training procedure that involves positional embedding adaptation and full parameter continuous pre-training. Furthermore, we propose compressed model variants trained via knowledge distillation. The models were evaluated on 25 tasks, including the KLEJ benchmark, a newly introduced financial task suite (FinBench), and other classification and regression tasks, specifically those requiring long-document understanding. The results demonstrate that our model achieves the best average performance among Polish and multilingual models, significantly outperforming competitive solutions in long-context tasks while maintaining comparable quality on short texts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_12191 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Long-Context Encoder Models for Polish Language Understanding Dadas, Sławomir Poświata, Rafał Kozłowski, Marek Grębowiec, Małgorzata Perełkiewicz, Michał Klimiuk, Paweł Boruta, Przemysław Computation and Language While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for discriminative tasks. However, classic encoders like BERT are limited by a short context window, which is insufficient for processing long documents. In this paper, we address this limitation for the Polish by introducing a high-quality Polish model capable of processing sequences of up to 8192 tokens. The model was developed by employing a two-stage training procedure that involves positional embedding adaptation and full parameter continuous pre-training. Furthermore, we propose compressed model variants trained via knowledge distillation. The models were evaluated on 25 tasks, including the KLEJ benchmark, a newly introduced financial task suite (FinBench), and other classification and regression tasks, specifically those requiring long-document understanding. The results demonstrate that our model achieves the best average performance among Polish and multilingual models, significantly outperforming competitive solutions in long-context tasks while maintaining comparable quality on short texts. |
| title | Long-Context Encoder Models for Polish Language Understanding |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2603.12191 |