Sailor: Open Language Models for South-East Asia

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dou, Longxu, Liu, Qian, Zeng, Guangtao, Guo, Jia, Zhou, Jiahui, Lu, Wei, Lin, Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916193469726720
author Dou, Longxu
Liu, Qian
Zeng, Guangtao
Guo, Jia
Zhou, Jiahui
Lu, Wei
Lin, Min
author_facet Dou, Longxu
Liu, Qian
Zeng, Guangtao
Guo, Jia
Zhou, Jiahui
Lu, Wei
Lin, Min
contents We present Sailor, a family of open language models ranging from 0.5B to 7B parameters, tailored for South-East Asian (SEA) languages. These models are continually pre-trained from Qwen1.5, a great language model for multilingual use cases. From Qwen1.5, Sailor models accept 200B to 400B tokens, primarily covering the languages of English, Chinese, Vietnamese, Thai, Indonesian, Malay, and Lao. The training leverages several techniques, including BPE dropout for improving the model robustness, aggressive data cleaning and deduplication, and small proxy models to optimize data mixture. Experimental results on four typical tasks indicate that Sailor models demonstrate strong performance across different benchmarks, including commonsense reasoning, question answering, reading comprehension and examination. Embracing the open-source spirit, we share our insights through this report to spark a wider interest in developing large language models for multilingual use cases.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sailor: Open Language Models for South-East Asia
Dou, Longxu
Liu, Qian
Zeng, Guangtao
Guo, Jia
Zhou, Jiahui
Lu, Wei
Lin, Min
Computation and Language
Artificial Intelligence
We present Sailor, a family of open language models ranging from 0.5B to 7B parameters, tailored for South-East Asian (SEA) languages. These models are continually pre-trained from Qwen1.5, a great language model for multilingual use cases. From Qwen1.5, Sailor models accept 200B to 400B tokens, primarily covering the languages of English, Chinese, Vietnamese, Thai, Indonesian, Malay, and Lao. The training leverages several techniques, including BPE dropout for improving the model robustness, aggressive data cleaning and deduplication, and small proxy models to optimize data mixture. Experimental results on four typical tasks indicate that Sailor models demonstrate strong performance across different benchmarks, including commonsense reasoning, question answering, reading comprehension and examination. Embracing the open-source spirit, we share our insights through this report to spark a wider interest in developing large language models for multilingual use cases.
title Sailor: Open Language Models for South-East Asia
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.03608