How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen-Yi, Andrea W, Jo, Unso Eun Seo, Lin, Lu Jia, Mimno, David
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911099655290880
author Wen-Yi, Andrea W
Jo, Unso Eun Seo
Lin, Lu Jia
Mimno, David
author_facet Wen-Yi, Andrea W
Jo, Unso Eun Seo
Lin, Lu Jia
Mimno, David
contents Contemporary language models are increasingly multilingual, but Chinese LLM developers must navigate complex political and business considerations of language diversity. Language policy in China aims at influencing the public discourse and governing a multi-ethnic society, and has gradually transitioned from a pluralist to a more assimilationist approach since 1949. We explore the impact of these influences on current language technology. We evaluate six open-source multilingual LLMs pre-trained by Chinese companies on 18 languages, spanning a wide range of Chinese, Asian, and Anglo-European languages. Our experiments show Chinese LLMs performance on diverse languages is indistinguishable from international LLMs. Similarly, the models' technical reports also show lack of consideration for pretraining data language coverage except for English and Mandarin Chinese. Examining Chinese AI policy, model experiments, and technical reports, we find no sign of any consistent policy, either for or against, language diversity in China's LLM development. This leaves a puzzling fact that while China regulates both the languages people use daily as well as language model development, they do not seem to have any policy on the languages in language models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_09652
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs
Wen-Yi, Andrea W
Jo, Unso Eun Seo
Lin, Lu Jia
Mimno, David
Computation and Language
Contemporary language models are increasingly multilingual, but Chinese LLM developers must navigate complex political and business considerations of language diversity. Language policy in China aims at influencing the public discourse and governing a multi-ethnic society, and has gradually transitioned from a pluralist to a more assimilationist approach since 1949. We explore the impact of these influences on current language technology. We evaluate six open-source multilingual LLMs pre-trained by Chinese companies on 18 languages, spanning a wide range of Chinese, Asian, and Anglo-European languages. Our experiments show Chinese LLMs performance on diverse languages is indistinguishable from international LLMs. Similarly, the models' technical reports also show lack of consideration for pretraining data language coverage except for English and Mandarin Chinese. Examining Chinese AI policy, model experiments, and technical reports, we find no sign of any consistent policy, either for or against, language diversity in China's LLM development. This leaves a puzzling fact that while China regulates both the languages people use daily as well as language model development, they do not seem to have any policy on the languages in language models.
title How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs
topic Computation and Language
url https://arxiv.org/abs/2407.09652