SeaLLMs -- Large Language Models for Southeast Asia
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917708496371712 |
|---|---|
| author | Nguyen, Xuan-Phi Zhang, Wenxuan Li, Xin Aljunied, Mahani Hu, Zhiqiang Shen, Chenhui Chia, Yew Ken Li, Xingxuan Wang, Jianyu Tan, Qingyu Cheng, Liying Chen, Guanzheng Deng, Yue Yang, Sen Liu, Chaoqun Zhang, Hang Bing, Lidong |
| author_facet | Nguyen, Xuan-Phi Zhang, Wenxuan Li, Xin Aljunied, Mahani Hu, Zhiqiang Shen, Chenhui Chia, Yew Ken Li, Xingxuan Wang, Jianyu Tan, Qingyu Cheng, Liying Chen, Guanzheng Deng, Yue Yang, Sen Liu, Chaoqun Zhang, Hang Bing, Lidong |
| contents | Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To address this imbalance, we introduce SeaLLMs, an innovative series of language models that specifically focuses on Southeast Asian (SEA) languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages. This allows them to respect and reflect local cultural norms, customs, stylistic preferences, and legal considerations. Our comprehensive evaluation demonstrates that SeaLLM-13b models exhibit superior performance across a wide spectrum of linguistic tasks and assistant-style instruction-following capabilities relative to comparable open-source models. Moreover, they outperform ChatGPT-3.5 in non-Latin languages, such as Thai, Khmer, Lao, and Burmese, by large margins while remaining lightweight and cost-effective to operate. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_00738 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | SeaLLMs -- Large Language Models for Southeast Asia Nguyen, Xuan-Phi Zhang, Wenxuan Li, Xin Aljunied, Mahani Hu, Zhiqiang Shen, Chenhui Chia, Yew Ken Li, Xingxuan Wang, Jianyu Tan, Qingyu Cheng, Liying Chen, Guanzheng Deng, Yue Yang, Sen Liu, Chaoqun Zhang, Hang Bing, Lidong Computation and Language Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To address this imbalance, we introduce SeaLLMs, an innovative series of language models that specifically focuses on Southeast Asian (SEA) languages. SeaLLMs are built upon the Llama-2 model and further advanced through continued pre-training with an extended vocabulary, specialized instruction and alignment tuning to better capture the intricacies of regional languages. This allows them to respect and reflect local cultural norms, customs, stylistic preferences, and legal considerations. Our comprehensive evaluation demonstrates that SeaLLM-13b models exhibit superior performance across a wide spectrum of linguistic tasks and assistant-style instruction-following capabilities relative to comparable open-source models. Moreover, they outperform ChatGPT-3.5 in non-Latin languages, such as Thai, Khmer, Lao, and Burmese, by large margins while remaining lightweight and cost-effective to operate. |
| title | SeaLLMs -- Large Language Models for Southeast Asia |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2312.00738 |