PersianMind: A Cross-Lingual Persian-English Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rostami, Pedram, Salemi, Ali, Dousti, Mohammad Javad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909071120007168
author Rostami, Pedram
Salemi, Ali
Dousti, Mohammad Javad
author_facet Rostami, Pedram
Salemi, Ali
Dousti, Mohammad Javad
contents Large language models demonstrate remarkable proficiency in various linguistic tasks and have extensive knowledge across various domains. Although they perform best in English, their ability in other languages is notable too. In contrast, open-source models, such as LLaMa, are primarily trained on English datasets, resulting in poor performance in non-English languages. In this paper, we introduce PersianMind, an open-source bilingual large language model which demonstrates comparable performance to closed-source GPT-3.5-turbo in the Persian language. By expanding LLaMa2's vocabulary with 10,000 Persian tokens and training it on a dataset comprising nearly 2 billion Persian tokens, we show that our approach preserves the model's English knowledge and employs transfer learning to excel at transferring task knowledge from one language to another.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06466
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PersianMind: A Cross-Lingual Persian-English Large Language Model
Rostami, Pedram
Salemi, Ali
Dousti, Mohammad Javad
Computation and Language
Artificial Intelligence
Large language models demonstrate remarkable proficiency in various linguistic tasks and have extensive knowledge across various domains. Although they perform best in English, their ability in other languages is notable too. In contrast, open-source models, such as LLaMa, are primarily trained on English datasets, resulting in poor performance in non-English languages. In this paper, we introduce PersianMind, an open-source bilingual large language model which demonstrates comparable performance to closed-source GPT-3.5-turbo in the Persian language. By expanding LLaMa2's vocabulary with 10,000 Persian tokens and training it on a dataset comprising nearly 2 billion Persian tokens, we show that our approach preserves the model's English knowledge and employs transfer learning to excel at transferring task knowledge from one language to another.
title PersianMind: A Cross-Lingual Persian-English Large Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.06466