Saved in:
Bibliographic Details
Main Authors: Kew, Tannon, Schottmann, Florian, Sennrich, Rico
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2312.12683
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916421040078848
author Kew, Tannon
Schottmann, Florian
Sennrich, Rico
author_facet Kew, Tannon
Schottmann, Florian
Sennrich, Rico
contents The vast majority of today's large language models (LLMs) are English-centric, having been pretrained predominantly on English text. Yet, in order to meet user expectations, models need to be able to respond appropriately in multiple languages once deployed in downstream applications. This requires strong cross-lingual transfer abilities. In this work, we investigate the minimal amount of multilinguality required during finetuning to elicit cross-lingual generalisation in English-centric LLMs. In experiments across four LLMs, we find that multilingual instruction tuning with as few as two to three languages is both necessary and sufficient to elicit effective cross-lingual generalisation, with the limiting factor being the degree to which a target language is seen during pretraining. Evaluations on five different tasks further reveal that multilingual instruction tuning is most beneficial for generative tasks that assume input/output language agreement, such as in chat settings, while being of less importance for highly structured classification-style tasks. Our code and data is available at https://github.com/ZurichNLP/multilingual-instruction-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12683
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
Kew, Tannon
Schottmann, Florian
Sennrich, Rico
Computation and Language
The vast majority of today's large language models (LLMs) are English-centric, having been pretrained predominantly on English text. Yet, in order to meet user expectations, models need to be able to respond appropriately in multiple languages once deployed in downstream applications. This requires strong cross-lingual transfer abilities. In this work, we investigate the minimal amount of multilinguality required during finetuning to elicit cross-lingual generalisation in English-centric LLMs. In experiments across four LLMs, we find that multilingual instruction tuning with as few as two to three languages is both necessary and sufficient to elicit effective cross-lingual generalisation, with the limiting factor being the degree to which a target language is seen during pretraining. Evaluations on five different tasks further reveal that multilingual instruction tuning is most beneficial for generative tasks that assume input/output language agreement, such as in chat settings, while being of less importance for highly structured classification-style tasks. Our code and data is available at https://github.com/ZurichNLP/multilingual-instruction-tuning.
title Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
topic Computation and Language
url https://arxiv.org/abs/2312.12683