Accelerating Multilingual Language Model for Excessively Tokenized Languages

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hong, Jimin, Lee, Gibbeum, Cho, Jaewoong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917741667024896
author Hong, Jimin
Lee, Gibbeum
Cho, Jaewoong
author_facet Hong, Jimin
Lee, Gibbeum
Cho, Jaewoong
contents Recent advancements in large language models (LLMs) have remarkably enhanced performances on a variety of tasks in multiple languages. However, tokenizers in LLMs trained primarily on English-centric corpora often overly fragment a text into character or Unicode-level tokens in non-Roman alphabetic languages, leading to inefficient text generation. We introduce a simple yet effective framework to accelerate text generation in such languages. Our approach involves employing a new language model head with a vocabulary set tailored to a specific target language for a pre-trained LLM. This is followed by fine-tuning the new head while incorporating a verification step to ensure the model's performance is preserved. We show that this targeted fine-tuning, while freezing other model parameters, effectively reduces token fragmentation for the target language. Our extensive experiments demonstrate that the proposed framework increases the generation speed by a factor of 1.7 while maintaining the performance of pre-trained multilingual models on target monolingual tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10660
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating Multilingual Language Model for Excessively Tokenized Languages
Hong, Jimin
Lee, Gibbeum
Cho, Jaewoong
Computation and Language
Artificial Intelligence
Recent advancements in large language models (LLMs) have remarkably enhanced performances on a variety of tasks in multiple languages. However, tokenizers in LLMs trained primarily on English-centric corpora often overly fragment a text into character or Unicode-level tokens in non-Roman alphabetic languages, leading to inefficient text generation. We introduce a simple yet effective framework to accelerate text generation in such languages. Our approach involves employing a new language model head with a vocabulary set tailored to a specific target language for a pre-trained LLM. This is followed by fine-tuning the new head while incorporating a verification step to ensure the model's performance is preserved. We show that this targeted fine-tuning, while freezing other model parameters, effectively reduces token fragmentation for the target language. Our extensive experiments demonstrate that the proposed framework increases the generation speed by a factor of 1.7 while maintaining the performance of pre-trained multilingual models on target monolingual tasks.
title Accelerating Multilingual Language Model for Excessively Tokenized Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.10660