An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamaguchi, Atsuki, Villavicencio, Aline, Aletras, Nikolaos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910620553576448
author Yamaguchi, Atsuki
Villavicencio, Aline
Aletras, Nikolaos
author_facet Yamaguchi, Atsuki
Villavicencio, Aline
Aletras, Nikolaos
contents The development of state-of-the-art generative large language models (LLMs) disproportionately relies on English-centric tokenizers, vocabulary and pre-training data. Despite the fact that some LLMs have multilingual capabilities, recent studies have shown that their inference efficiency deteriorates when generating text in languages other than English. This results in increased inference time and costs. Cross-lingual vocabulary adaptation (CVA) methods have been proposed for adapting models to a target language aiming to improve downstream performance. However, the effectiveness of these methods on increasing inference efficiency of generative LLMs has yet to be explored. In this paper, we perform an empirical study of five CVA methods on four generative LLMs (including monolingual and multilingual models) across four typologically-diverse languages and four natural language understanding tasks. We find that CVA substantially contributes to LLM inference speedups of up to 271.5\%. We also show that adapting LLMs that have been pre-trained on more balanced multilingual data results in downstream performance comparable to the original models.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10712
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
Yamaguchi, Atsuki
Villavicencio, Aline
Aletras, Nikolaos
Computation and Language
Artificial Intelligence
The development of state-of-the-art generative large language models (LLMs) disproportionately relies on English-centric tokenizers, vocabulary and pre-training data. Despite the fact that some LLMs have multilingual capabilities, recent studies have shown that their inference efficiency deteriorates when generating text in languages other than English. This results in increased inference time and costs. Cross-lingual vocabulary adaptation (CVA) methods have been proposed for adapting models to a target language aiming to improve downstream performance. However, the effectiveness of these methods on increasing inference efficiency of generative LLMs has yet to be explored. In this paper, we perform an empirical study of five CVA methods on four generative LLMs (including monolingual and multilingual models) across four typologically-diverse languages and four natural language understanding tasks. We find that CVA substantially contributes to LLM inference speedups of up to 271.5\%. We also show that adapting LLMs that have been pre-trained on more balanced multilingual data results in downstream performance comparable to the original models.
title An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.10712