Long-Context Encoder Models for Polish Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dadas, Sławomir, Poświata, Rafał, Kozłowski, Marek, Grębowiec, Małgorzata, Perełkiewicz, Michał, Klimiuk, Paweł, Boruta, Przemysław
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918384487104512
author Dadas, Sławomir
Poświata, Rafał
Kozłowski, Marek
Grębowiec, Małgorzata
Perełkiewicz, Michał
Klimiuk, Paweł
Boruta, Przemysław
author_facet Dadas, Sławomir
Poświata, Rafał
Kozłowski, Marek
Grębowiec, Małgorzata
Perełkiewicz, Michał
Klimiuk, Paweł
Boruta, Przemysław
contents While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for discriminative tasks. However, classic encoders like BERT are limited by a short context window, which is insufficient for processing long documents. In this paper, we address this limitation for the Polish by introducing a high-quality Polish model capable of processing sequences of up to 8192 tokens. The model was developed by employing a two-stage training procedure that involves positional embedding adaptation and full parameter continuous pre-training. Furthermore, we propose compressed model variants trained via knowledge distillation. The models were evaluated on 25 tasks, including the KLEJ benchmark, a newly introduced financial task suite (FinBench), and other classification and regression tasks, specifically those requiring long-document understanding. The results demonstrate that our model achieves the best average performance among Polish and multilingual models, significantly outperforming competitive solutions in long-context tasks while maintaining comparable quality on short texts.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12191
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Long-Context Encoder Models for Polish Language Understanding
Dadas, Sławomir
Poświata, Rafał
Kozłowski, Marek
Grębowiec, Małgorzata
Perełkiewicz, Michał
Klimiuk, Paweł
Boruta, Przemysław
Computation and Language
While decoder-only Large Language Models (LLMs) have recently dominated the NLP landscape, encoder-only architectures remain a cost-effective and parameter-efficient standard for discriminative tasks. However, classic encoders like BERT are limited by a short context window, which is insufficient for processing long documents. In this paper, we address this limitation for the Polish by introducing a high-quality Polish model capable of processing sequences of up to 8192 tokens. The model was developed by employing a two-stage training procedure that involves positional embedding adaptation and full parameter continuous pre-training. Furthermore, we propose compressed model variants trained via knowledge distillation. The models were evaluated on 25 tasks, including the KLEJ benchmark, a newly introduced financial task suite (FinBench), and other classification and regression tasks, specifically those requiring long-document understanding. The results demonstrate that our model achieves the best average performance among Polish and multilingual models, significantly outperforming competitive solutions in long-context tasks while maintaining comparable quality on short texts.
title Long-Context Encoder Models for Polish Language Understanding
topic Computation and Language
url https://arxiv.org/abs/2603.12191