Sailor: Open Language Models for South-East Asia
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916193469726720 |
|---|---|
| author | Dou, Longxu Liu, Qian Zeng, Guangtao Guo, Jia Zhou, Jiahui Lu, Wei Lin, Min |
| author_facet | Dou, Longxu Liu, Qian Zeng, Guangtao Guo, Jia Zhou, Jiahui Lu, Wei Lin, Min |
| contents | We present Sailor, a family of open language models ranging from 0.5B to 7B parameters, tailored for South-East Asian (SEA) languages. These models are continually pre-trained from Qwen1.5, a great language model for multilingual use cases. From Qwen1.5, Sailor models accept 200B to 400B tokens, primarily covering the languages of English, Chinese, Vietnamese, Thai, Indonesian, Malay, and Lao. The training leverages several techniques, including BPE dropout for improving the model robustness, aggressive data cleaning and deduplication, and small proxy models to optimize data mixture. Experimental results on four typical tasks indicate that Sailor models demonstrate strong performance across different benchmarks, including commonsense reasoning, question answering, reading comprehension and examination. Embracing the open-source spirit, we share our insights through this report to spark a wider interest in developing large language models for multilingual use cases. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_03608 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Sailor: Open Language Models for South-East Asia Dou, Longxu Liu, Qian Zeng, Guangtao Guo, Jia Zhou, Jiahui Lu, Wei Lin, Min Computation and Language Artificial Intelligence We present Sailor, a family of open language models ranging from 0.5B to 7B parameters, tailored for South-East Asian (SEA) languages. These models are continually pre-trained from Qwen1.5, a great language model for multilingual use cases. From Qwen1.5, Sailor models accept 200B to 400B tokens, primarily covering the languages of English, Chinese, Vietnamese, Thai, Indonesian, Malay, and Lao. The training leverages several techniques, including BPE dropout for improving the model robustness, aggressive data cleaning and deduplication, and small proxy models to optimize data mixture. Experimental results on four typical tasks indicate that Sailor models demonstrate strong performance across different benchmarks, including commonsense reasoning, question answering, reading comprehension and examination. Embracing the open-source spirit, we share our insights through this report to spark a wider interest in developing large language models for multilingual use cases. |
| title | Sailor: Open Language Models for South-East Asia |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2404.03608 |