Low-Resource Machine Translation through the Lens of Personalized Federated Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Moskvoretskii, Viktor, Tupitsa, Nazarii, Biemann, Chris, Horváth, Samuel, Gorbunov, Eduard, Nikishina, Irina
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913620224376832
author Moskvoretskii, Viktor
Tupitsa, Nazarii
Biemann, Chris
Horváth, Samuel
Gorbunov, Eduard
Nikishina, Irina
author_facet Moskvoretskii, Viktor
Tupitsa, Nazarii
Biemann, Chris
Horváth, Samuel
Gorbunov, Eduard
Nikishina, Irina
contents We present a new approach called MeritOpt based on the Personalized Federated Learning algorithm MeritFed that can be applied to Natural Language Tasks with heterogeneous data. We evaluate it on the Low-Resource Machine Translation task, using the datasets of South East Asian and Finno-Ugric languages. In addition to its effectiveness, MeritOpt is also highly interpretable, as it can be applied to track the impact of each language used for training. Our analysis reveals that target dataset size affects weight distribution across auxiliary languages, that unrelated languages do not interfere with the training, and auxiliary optimizer parameters have minimal impact. Our approach is easy to apply with a few lines of code, and we provide scripts for reproducing the experiments at https://github.com/VityaVitalich/MeritOpt.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12564
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-Resource Machine Translation through the Lens of Personalized Federated Learning
Moskvoretskii, Viktor
Tupitsa, Nazarii
Biemann, Chris
Horváth, Samuel
Gorbunov, Eduard
Nikishina, Irina
Computation and Language
Machine Learning
We present a new approach called MeritOpt based on the Personalized Federated Learning algorithm MeritFed that can be applied to Natural Language Tasks with heterogeneous data. We evaluate it on the Low-Resource Machine Translation task, using the datasets of South East Asian and Finno-Ugric languages. In addition to its effectiveness, MeritOpt is also highly interpretable, as it can be applied to track the impact of each language used for training. Our analysis reveals that target dataset size affects weight distribution across auxiliary languages, that unrelated languages do not interfere with the training, and auxiliary optimizer parameters have minimal impact. Our approach is easy to apply with a few lines of code, and we provide scripts for reproducing the experiments at https://github.com/VityaVitalich/MeritOpt.
title Low-Resource Machine Translation through the Lens of Personalized Federated Learning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.12564