Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Glocker, Kevin, Kukk, Kätriin, Oji, Romina, Bollmann, Marcel, Kuhlmann, Marco, Kunz, Jenny
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911314182406144
author Glocker, Kevin
Kukk, Kätriin
Oji, Romina
Bollmann, Marcel
Kuhlmann, Marco
Kunz, Jenny
author_facet Glocker, Kevin
Kukk, Kätriin
Oji, Romina
Bollmann, Marcel
Kuhlmann, Marco
Kunz, Jenny
contents Achieving high-performing language models which include medium- and lower-resource languages remains a challenge. Massively multilingual models still underperform compared to language-specific adaptations, especially at smaller model scales. In this work, we investigate scaling as an efficient strategy for adapting pretrained models to new target languages. Through comprehensive scaling ablations with approximately FLOP-matched models, we test whether upscaling an English base model enables more effective and resource-efficient adaptation than standard continued pretraining. We find that, once exposed to sufficient target-language data, larger upscaled models can match or surpass the performance of smaller models continually pretrained on much more data, demonstrating the benefits of scaling for data efficiency. Scaling also helps preserve the base model's capabilities in English, thus reducing catastrophic forgetting. Finally, we explore whether such scaled, language-specific models can be merged to construct modular and flexible multilingual systems. We find that while merging remains less effective than joint multilingual training, upscaled merges perform better than smaller ones. We observe large performance differences across merging methods, suggesting potential for improvement through merging approaches specialized for language-level integration.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
Glocker, Kevin
Kukk, Kätriin
Oji, Romina
Bollmann, Marcel
Kuhlmann, Marco
Kunz, Jenny
Computation and Language
Artificial Intelligence
Machine Learning
Achieving high-performing language models which include medium- and lower-resource languages remains a challenge. Massively multilingual models still underperform compared to language-specific adaptations, especially at smaller model scales. In this work, we investigate scaling as an efficient strategy for adapting pretrained models to new target languages. Through comprehensive scaling ablations with approximately FLOP-matched models, we test whether upscaling an English base model enables more effective and resource-efficient adaptation than standard continued pretraining. We find that, once exposed to sufficient target-language data, larger upscaled models can match or surpass the performance of smaller models continually pretrained on much more data, demonstrating the benefits of scaling for data efficiency. Scaling also helps preserve the base model's capabilities in English, thus reducing catastrophic forgetting. Finally, we explore whether such scaled, language-specific models can be merged to construct modular and flexible multilingual systems. We find that while merging remains less effective than joint multilingual training, upscaled merges perform better than smaller ones. We observe large performance differences across merging methods, suggesting potential for improvement through merging approaches specialized for language-level integration.
title Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.10772