Pretraining Finnish ModernBERTs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911261767237632 |
|---|---|
| author | Reunamo, Akseli Peltonen, Laura-Maria Moen, Hans Pyysalo, Sampo |
| author_facet | Reunamo, Akseli Peltonen, Laura-Maria Moen, Hans Pyysalo, Sampo |
| contents | This paper reports on pretraining ModernBERT encoder models in six different sizes, ranging from 51M to 475M parameters, with a focus on limited multilingualism, emphasizing languages relevant to Finland. Our models are competitive with, or superior to, existing multilingual models. They outperform monolingual models on tasks that require a context longer than 512 tokens. We present empirical results on using different data in the final stage of training. The code and models are publicly released. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_09213 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Pretraining Finnish ModernBERTs Reunamo, Akseli Peltonen, Laura-Maria Moen, Hans Pyysalo, Sampo Computation and Language This paper reports on pretraining ModernBERT encoder models in six different sizes, ranging from 51M to 475M parameters, with a focus on limited multilingualism, emphasizing languages relevant to Finland. Our models are competitive with, or superior to, existing multilingual models. They outperform monolingual models on tasks that require a context longer than 512 tokens. We present empirical results on using different data in the final stage of training. The code and models are publicly released. |
| title | Pretraining Finnish ModernBERTs |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2511.09213 |