On Multilingual Encoder Language Model Compression for Low-Resource Languages

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gurgurov, Daniil, Gregor, Michal, van Genabith, Josef, Ostermann, Simon
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912689405558784
author Gurgurov, Daniil
Gregor, Michal
van Genabith, Josef
Ostermann, Simon
author_facet Gurgurov, Daniil
Gregor, Michal
van Genabith, Josef
Ostermann, Simon
contents In this paper, we combine two-step knowledge distillation, structured pruning, truncation, and vocabulary trimming for extremely compressing multilingual encoder-only language models for low-resource languages. Our novel approach systematically combines existing techniques and takes them to the extreme, reducing layer depth, feed-forward hidden size, and intermediate layer embedding size to create significantly smaller monolingual models while retaining essential language-specific knowledge. We achieve compression rates of up to 92% while maintaining competitive performance, with average drops of 2-10% for moderate compression and 8-13% at maximum compression in four downstream tasks, including sentiment analysis, topic classification, named entity recognition, and part-of-speech tagging, across three low-resource languages. Notably, the performance degradation correlates with the amount of language-specific data in the teacher model, with larger datasets resulting in smaller performance losses. Additionally, we conduct ablation studies to identify the best practices for multilingual model compression using these techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16956
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Multilingual Encoder Language Model Compression for Low-Resource Languages
Gurgurov, Daniil
Gregor, Michal
van Genabith, Josef
Ostermann, Simon
Computation and Language
In this paper, we combine two-step knowledge distillation, structured pruning, truncation, and vocabulary trimming for extremely compressing multilingual encoder-only language models for low-resource languages. Our novel approach systematically combines existing techniques and takes them to the extreme, reducing layer depth, feed-forward hidden size, and intermediate layer embedding size to create significantly smaller monolingual models while retaining essential language-specific knowledge. We achieve compression rates of up to 92% while maintaining competitive performance, with average drops of 2-10% for moderate compression and 8-13% at maximum compression in four downstream tasks, including sentiment analysis, topic classification, named entity recognition, and part-of-speech tagging, across three low-resource languages. Notably, the performance degradation correlates with the amount of language-specific data in the teacher model, with larger datasets resulting in smaller performance losses. Additionally, we conduct ablation studies to identify the best practices for multilingual model compression using these techniques.
title On Multilingual Encoder Language Model Compression for Low-Resource Languages
topic Computation and Language
url https://arxiv.org/abs/2505.16956