How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamaguchi, Atsuki, Villavicencio, Aline, Aletras, Nikolaos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909929172893696
author Yamaguchi, Atsuki
Villavicencio, Aline
Aletras, Nikolaos
author_facet Yamaguchi, Atsuki
Villavicencio, Aline
Aletras, Nikolaos
contents Large language models (LLMs) have shown remarkable capabilities in many languages beyond English. Yet, LLMs require more inference steps when generating non-English text due to their reliance on English-centric tokenizers and vocabulary, resulting in higher usage costs to non-English speakers. Vocabulary expansion with target language tokens is a widely used cross-lingual vocabulary adaptation approach to remedy this issue. Despite its effectiveness in inference speedup, previous work on vocabulary expansion has focused on high-resource settings assuming access to a substantial amount of target language data to effectively initialize the embeddings of the new tokens and adapt the LLM to the target language. However, vocabulary expansion in low-resource settings has yet to be explored. In this article, we investigate vocabulary expansion in low-resource settings by considering embedding initialization methods and continual pre-training strategies. Through extensive experiments across typologically diverse languages, tasks and models, we establish a set of strategies to perform vocabulary expansion for faster inference, while striving to maintain competitive downstream performance to baselines. This is achieved with only 30K sentences ($\sim$0.01GB text data) from the target language.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11477
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
Yamaguchi, Atsuki
Villavicencio, Aline
Aletras, Nikolaos
Computation and Language
Artificial Intelligence
Large language models (LLMs) have shown remarkable capabilities in many languages beyond English. Yet, LLMs require more inference steps when generating non-English text due to their reliance on English-centric tokenizers and vocabulary, resulting in higher usage costs to non-English speakers. Vocabulary expansion with target language tokens is a widely used cross-lingual vocabulary adaptation approach to remedy this issue. Despite its effectiveness in inference speedup, previous work on vocabulary expansion has focused on high-resource settings assuming access to a substantial amount of target language data to effectively initialize the embeddings of the new tokens and adapt the LLM to the target language. However, vocabulary expansion in low-resource settings has yet to be explored. In this article, we investigate vocabulary expansion in low-resource settings by considering embedding initialization methods and continual pre-training strategies. Through extensive experiments across typologically diverse languages, tasks and models, we establish a set of strategies to perform vocabulary expansion for faster inference, while striving to maintain competitive downstream performance to baselines. This is achieved with only 30K sentences ($\sim$0.01GB text data) from the target language.
title How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.11477