How do Large Language Models Handle Multilingualism?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yiran, Zhang, Wenxuan, Chen, Guizhen, Kawaguchi, Kenji, Bing, Lidong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913574767558656
author Zhao, Yiran
Zhang, Wenxuan
Chen, Guizhen
Kawaguchi, Kenji
Bing, Lidong
author_facet Zhao, Yiran
Zhang, Wenxuan
Chen, Guizhen
Kawaguchi, Kenji
Bing, Lidong
contents Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relationships between network structures and certain capabilities, we hypothesize the LLM's multilingual workflow ($\texttt{MWork}$): LLMs initially understand the query, converting multilingual inputs into English for task-solving. In the intermediate layers, they employ English for thinking and incorporate multilingual knowledge with self-attention and feed-forward structures, respectively. In the final layers, LLMs generate responses aligned with the original language of the query. To verify $\texttt{MWork}$, we introduce Parallel Language-specific Neuron Detection ($\texttt{PLND}$) to identify activated neurons for inputs in different languages without any labeled data. Using $\texttt{PLND}$, we validate $\texttt{MWork}$ through extensive experiments involving the deactivation of language-specific neurons across various layers and structures. Moreover, $\texttt{MWork}$ allows fine-tuning of language-specific neurons with a small dataset, enhancing multilingual abilities in a specific language without compromising others. This approach results in an average improvement of $3.6\%$ for high-resource languages and $2.3\%$ for low-resource languages across all tasks with just $400$ documents.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18815
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How do Large Language Models Handle Multilingualism?
Zhao, Yiran
Zhang, Wenxuan
Chen, Guizhen
Kawaguchi, Kenji
Bing, Lidong
Computation and Language
Artificial Intelligence
Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relationships between network structures and certain capabilities, we hypothesize the LLM's multilingual workflow ($\texttt{MWork}$): LLMs initially understand the query, converting multilingual inputs into English for task-solving. In the intermediate layers, they employ English for thinking and incorporate multilingual knowledge with self-attention and feed-forward structures, respectively. In the final layers, LLMs generate responses aligned with the original language of the query. To verify $\texttt{MWork}$, we introduce Parallel Language-specific Neuron Detection ($\texttt{PLND}$) to identify activated neurons for inputs in different languages without any labeled data. Using $\texttt{PLND}$, we validate $\texttt{MWork}$ through extensive experiments involving the deactivation of language-specific neurons across various layers and structures. Moreover, $\texttt{MWork}$ allows fine-tuning of language-specific neurons with a small dataset, enhancing multilingual abilities in a specific language without compromising others. This approach results in an average improvement of $3.6\%$ for high-resource languages and $2.3\%$ for low-resource languages across all tasks with just $400$ documents.
title How do Large Language Models Handle Multilingualism?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.18815