MaLA-500: Massive Language Adaptation of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Peiqin, Ji, Shaoxiong, Tiedemann, Jörg, Martins, André F. T., Schütze, Hinrich
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916190688903168
author Lin, Peiqin
Ji, Shaoxiong
Tiedemann, Jörg
Martins, André F. T.
Schütze, Hinrich
author_facet Lin, Peiqin
Ji, Shaoxiong
Tiedemann, Jörg
Martins, André F. T.
Schütze, Hinrich
contents Large language models (LLMs) have advanced the state of the art in natural language processing. However, their predominant design for English or a limited set of languages creates a substantial gap in their effectiveness for low-resource languages. To bridge this gap, we introduce MaLA-500, a novel large language model designed to cover an extensive range of 534 languages. To train MaLA-500, we employ vocabulary extension and continued pretraining on LLaMA 2 with Glot500-c. Our intrinsic evaluation demonstrates that MaLA-500 is better at predicting the given texts of low-resource languages than existing multilingual LLMs. Moreover, the extrinsic evaluation of in-context learning shows that MaLA-500 outperforms previous LLMs on SIB200 and Taxi1500 by a significant margin, i.e., 11.68% and 4.82% marco-average accuracy across languages. We release MaLA-500 at https://huggingface.co/MaLA-LM
format Preprint
id arxiv_https___arxiv_org_abs_2401_13303
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MaLA-500: Massive Language Adaptation of Large Language Models
Lin, Peiqin
Ji, Shaoxiong
Tiedemann, Jörg
Martins, André F. T.
Schütze, Hinrich
Computation and Language
Large language models (LLMs) have advanced the state of the art in natural language processing. However, their predominant design for English or a limited set of languages creates a substantial gap in their effectiveness for low-resource languages. To bridge this gap, we introduce MaLA-500, a novel large language model designed to cover an extensive range of 534 languages. To train MaLA-500, we employ vocabulary extension and continued pretraining on LLaMA 2 with Glot500-c. Our intrinsic evaluation demonstrates that MaLA-500 is better at predicting the given texts of low-resource languages than existing multilingual LLMs. Moreover, the extrinsic evaluation of in-context learning shows that MaLA-500 outperforms previous LLMs on SIB200 and Taxi1500 by a significant margin, i.e., 11.68% and 4.82% marco-average accuracy across languages. We release MaLA-500 at https://huggingface.co/MaLA-LM
title MaLA-500: Massive Language Adaptation of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2401.13303