Low-Resource Machine Translation through the Lens of Personalized Federated Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913620224376832 |
|---|---|
| author | Moskvoretskii, Viktor Tupitsa, Nazarii Biemann, Chris Horváth, Samuel Gorbunov, Eduard Nikishina, Irina |
| author_facet | Moskvoretskii, Viktor Tupitsa, Nazarii Biemann, Chris Horváth, Samuel Gorbunov, Eduard Nikishina, Irina |
| contents | We present a new approach called MeritOpt based on the Personalized Federated Learning algorithm MeritFed that can be applied to Natural Language Tasks with heterogeneous data. We evaluate it on the Low-Resource Machine Translation task, using the datasets of South East Asian and Finno-Ugric languages. In addition to its effectiveness, MeritOpt is also highly interpretable, as it can be applied to track the impact of each language used for training. Our analysis reveals that target dataset size affects weight distribution across auxiliary languages, that unrelated languages do not interfere with the training, and auxiliary optimizer parameters have minimal impact. Our approach is easy to apply with a few lines of code, and we provide scripts for reproducing the experiments at https://github.com/VityaVitalich/MeritOpt. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_12564 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Low-Resource Machine Translation through the Lens of Personalized Federated Learning Moskvoretskii, Viktor Tupitsa, Nazarii Biemann, Chris Horváth, Samuel Gorbunov, Eduard Nikishina, Irina Computation and Language Machine Learning We present a new approach called MeritOpt based on the Personalized Federated Learning algorithm MeritFed that can be applied to Natural Language Tasks with heterogeneous data. We evaluate it on the Low-Resource Machine Translation task, using the datasets of South East Asian and Finno-Ugric languages. In addition to its effectiveness, MeritOpt is also highly interpretable, as it can be applied to track the impact of each language used for training. Our analysis reveals that target dataset size affects weight distribution across auxiliary languages, that unrelated languages do not interfere with the training, and auxiliary optimizer parameters have minimal impact. Our approach is easy to apply with a few lines of code, and we provide scripts for reproducing the experiments at https://github.com/VityaVitalich/MeritOpt. |
| title | Low-Resource Machine Translation through the Lens of Personalized Federated Learning |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2406.12564 |