Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sani, Samin Mahdizadeh, Sadeghi, Pouya, Vu, Thuy-Trang, Yaghoobzadeh, Yadollah, Haffari, Gholamreza
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917887258656768
author Sani, Samin Mahdizadeh
Sadeghi, Pouya
Vu, Thuy-Trang
Yaghoobzadeh, Yadollah
Haffari, Gholamreza
author_facet Sani, Samin Mahdizadeh
Sadeghi, Pouya
Vu, Thuy-Trang
Yaghoobzadeh, Yadollah
Haffari, Gholamreza
contents Large language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-resource languages. In this study, we explore adding a new language, i.e., Persian, to Llama (a model with a limited understanding of Persian) using parameter-efficient fine-tuning. We employ a multi-stage approach involving pretraining on monolingual Persian data, aligning representations through bilingual pretraining and instruction datasets, and instruction-tuning with task-specific datasets. We evaluate the model's performance at each stage on generation and classification tasks. Our findings suggest that incorporating the Persian language, through bilingual data alignment, can enhance classification accuracy for Persian tasks, with no adverse impact and sometimes even improvements on English tasks. Additionally, the results highlight the model's initial strength as a critical factor when working with limited training data, with cross-lingual alignment offering minimal benefits for the low-resource language. Knowledge transfer from English to Persian has a marginal effect, primarily benefiting simple classification tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13375
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
Sani, Samin Mahdizadeh
Sadeghi, Pouya
Vu, Thuy-Trang
Yaghoobzadeh, Yadollah
Haffari, Gholamreza
Computation and Language
Large language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-resource languages. In this study, we explore adding a new language, i.e., Persian, to Llama (a model with a limited understanding of Persian) using parameter-efficient fine-tuning. We employ a multi-stage approach involving pretraining on monolingual Persian data, aligning representations through bilingual pretraining and instruction datasets, and instruction-tuning with task-specific datasets. We evaluate the model's performance at each stage on generation and classification tasks. Our findings suggest that incorporating the Persian language, through bilingual data alignment, can enhance classification accuracy for Persian tasks, with no adverse impact and sometimes even improvements on English tasks. Additionally, the results highlight the model's initial strength as a critical factor when working with limited training data, with cross-lingual alignment offering minimal benefits for the low-resource language. Knowledge transfer from English to Persian has a marginal effect, primarily benefiting simple classification tasks.
title Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
topic Computation and Language
url https://arxiv.org/abs/2412.13375