SeaLLMs -- Large Language Models for Southeast Asia

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Xuan-Phi, Zhang, Wenxuan, Li, Xin, Aljunied, Mahani, Hu, Zhiqiang, Shen, Chenhui, Chia, Yew Ken, Li, Xingxuan, Wang, Jianyu, Tan, Qingyu, Cheng, Liying, Chen, Guanzheng, Deng, Yue, Yang, Sen, Liu, Chaoqun, Zhang, Hang, Bing, Lidong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917708496371712
author Nguyen, Xuan-Phi
Zhang, Wenxuan
Li, Xin
Aljunied, Mahani
Hu, Zhiqiang
Shen, Chenhui
Chia, Yew Ken
Li, Xingxuan
Wang, Jianyu
Tan, Qingyu
Cheng, Liying
Chen, Guanzheng
Deng, Yue
Yang, Sen
Liu, Chaoqun
Zhang, Hang
Bing, Lidong
author_facet Nguyen, Xuan-Phi
Zhang, Wenxuan
Li, Xin
Aljunied, Mahani
Hu, Zhiqiang
Shen, Chenhui
Chia, Yew Ken
Li, Xingxuan
Wang, Jianyu
Tan, Qingyu
Cheng, Liying
Chen, Guanzheng
Deng, Yue
Yang, Sen
Liu, Chaoqun
Zhang, Hang
Bing, Lidong
contents Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To address this imbalance, we introduce SeaLLMs, an innovative series of language models that specifically focuses on Southeast Asian (SEA) languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages. This allows them to respect and reflect local cultural norms, customs, stylistic preferences, and legal considerations. Our comprehensive evaluation demonstrates that SeaLLM-13b models exhibit superior performance across a wide spectrum of linguistic tasks and assistant-style instruction-following capabilities relative to comparable open-source models. Moreover, they outperform ChatGPT-3.5 in non-Latin languages, such as Thai, Khmer, Lao, and Burmese, by large margins while remaining lightweight and cost-effective to operate.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00738
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SeaLLMs -- Large Language Models for Southeast Asia
Nguyen, Xuan-Phi
Zhang, Wenxuan
Li, Xin
Aljunied, Mahani
Hu, Zhiqiang
Shen, Chenhui
Chia, Yew Ken
Li, Xingxuan
Wang, Jianyu
Tan, Qingyu
Cheng, Liying
Chen, Guanzheng
Deng, Yue
Yang, Sen
Liu, Chaoqun
Zhang, Hang
Bing, Lidong
Computation and Language
Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To address this imbalance, we introduce SeaLLMs, an innovative series of language models that specifically focuses on Southeast Asian (SEA) languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages. This allows them to respect and reflect local cultural norms, customs, stylistic preferences, and legal considerations. Our comprehensive evaluation demonstrates that SeaLLM-13b models exhibit superior performance across a wide spectrum of linguistic tasks and assistant-style instruction-following capabilities relative to comparable open-source models. Moreover, they outperform ChatGPT-3.5 in non-Latin languages, such as Thai, Khmer, Lao, and Burmese, by large margins while remaining lightweight and cost-effective to operate.
title SeaLLMs -- Large Language Models for Southeast Asia
topic Computation and Language
url https://arxiv.org/abs/2312.00738