Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chirkova, Nadezhda, Liang, Sheng, Nikoulina, Vassilina
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917646219345920
author Chirkova, Nadezhda
Liang, Sheng
Nikoulina, Vassilina
author_facet Chirkova, Nadezhda
Liang, Sheng
Nikoulina, Vassilina
contents Zero-shot cross-lingual knowledge transfer enables the multilingual pretrained language model (mPLM), finetuned on a task in one language, make predictions for this task in other languages. While being broadly studied for natural language understanding tasks, the described setting is understudied for generation. Previous works notice a frequent problem of generation in a wrong language and propose approaches to address it, usually using mT5 as a backbone model. In this work, we test alternative mPLMs, such as mBART and NLLB-200, considering full finetuning and parameter-efficient finetuning with adapters. We find that mBART with adapters performs similarly to mT5 of the same size, and NLLB-200 can be competitive in some cases. We also underline the importance of tuning learning rate used for finetuning, which helps to alleviate the problem of generation in the wrong language.
format Preprint
id arxiv_https___arxiv_org_abs_2310_09917
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
Chirkova, Nadezhda
Liang, Sheng
Nikoulina, Vassilina
Computation and Language
Zero-shot cross-lingual knowledge transfer enables the multilingual pretrained language model (mPLM), finetuned on a task in one language, make predictions for this task in other languages. While being broadly studied for natural language understanding tasks, the described setting is understudied for generation. Previous works notice a frequent problem of generation in a wrong language and propose approaches to address it, usually using mT5 as a backbone model. In this work, we test alternative mPLMs, such as mBART and NLLB-200, considering full finetuning and parameter-efficient finetuning with adapters. We find that mBART with adapters performs similarly to mT5 of the same size, and NLLB-200 can be competitive in some cases. We also underline the importance of tuning learning rate used for finetuning, which helps to alleviate the problem of generation in the wrong language.
title Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
topic Computation and Language
url https://arxiv.org/abs/2310.09917