Pretraining Finnish ModernBERTs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reunamo, Akseli, Peltonen, Laura-Maria, Moen, Hans, Pyysalo, Sampo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911261767237632
author Reunamo, Akseli
Peltonen, Laura-Maria
Moen, Hans
Pyysalo, Sampo
author_facet Reunamo, Akseli
Peltonen, Laura-Maria
Moen, Hans
Pyysalo, Sampo
contents This paper reports on pretraining ModernBERT encoder models in six different sizes, ranging from 51M to 475M parameters, with a focus on limited multilingualism, emphasizing languages relevant to Finland. Our models are competitive with, or superior to, existing multilingual models. They outperform monolingual models on tasks that require a context longer than 512 tokens. We present empirical results on using different data in the final stage of training. The code and models are publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pretraining Finnish ModernBERTs
Reunamo, Akseli
Peltonen, Laura-Maria
Moen, Hans
Pyysalo, Sampo
Computation and Language
This paper reports on pretraining ModernBERT encoder models in six different sizes, ranging from 51M to 475M parameters, with a focus on limited multilingualism, emphasizing languages relevant to Finland. Our models are competitive with, or superior to, existing multilingual models. They outperform monolingual models on tasks that require a context longer than 512 tokens. We present empirical results on using different data in the final stage of training. The code and models are publicly released.
title Pretraining Finnish ModernBERTs
topic Computation and Language
url https://arxiv.org/abs/2511.09213