BERnaT: Basque Encoders for Representing Natural Textual Diversity

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Azurmendi, Ekhi, de Landa, Joseba Fernandez, Bengoetxea, Jaione, Heredia, Maite, Etxaniz, Julen, Zubillaga, Mikel, Soraluze, Ander, Soroa, Aitor
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911536905191424
author Azurmendi, Ekhi
de Landa, Joseba Fernandez
Bengoetxea, Jaione
Heredia, Maite
Etxaniz, Julen
Zubillaga, Mikel
Soraluze, Ander
Soroa, Aitor
author_facet Azurmendi, Ekhi
de Landa, Joseba Fernandez
Bengoetxea, Jaione
Heredia, Maite
Etxaniz, Julen
Zubillaga, Mikel
Soraluze, Ander
Soroa, Aitor
contents Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal, historical, informal, etc.) rather than relying solely on standardized text. Focusing on the Basque language, we construct new corpora combining standard, social media, and historical sources, and pre-train the BERnaT family of encoder-only models in three configurations: standard, diverse, and combined. We further propose an evaluation framework that separates Natural Language Understanding (NLU) tasks into standard and diverse subsets to assess linguistic generalization. Results show that models trained on both standard and diverse data consistently outperform those trained on standard corpora, improving performance across all task types without compromising standard benchmark accuracy. These findings highlight the importance of linguistic diversity in building inclusive, generalizable language models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03903
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BERnaT: Basque Encoders for Representing Natural Textual Diversity
Azurmendi, Ekhi
de Landa, Joseba Fernandez
Bengoetxea, Jaione
Heredia, Maite
Etxaniz, Julen
Zubillaga, Mikel
Soraluze, Ander
Soroa, Aitor
Computation and Language
Artificial Intelligence
Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational biases. In this paper, we argue that language models should aim to capture the full spectrum of language variation (dialectal, historical, informal, etc.) rather than relying solely on standardized text. Focusing on the Basque language, we construct new corpora combining standard, social media, and historical sources, and pre-train the BERnaT family of encoder-only models in three configurations: standard, diverse, and combined. We further propose an evaluation framework that separates Natural Language Understanding (NLU) tasks into standard and diverse subsets to assess linguistic generalization. Results show that models trained on both standard and diverse data consistently outperform those trained on standard corpora, improving performance across all task types without compromising standard benchmark accuracy. These findings highlight the importance of linguistic diversity in building inclusive, generalizable language models.
title BERnaT: Basque Encoders for Representing Natural Textual Diversity
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.03903