Lossless Vocabulary Reduction for Auto-Regressive Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chijiwa, Daiki, Hasegawa, Taku, Nishida, Kyosuke, Yamaguchi, Shin'ya, Ohba, Tomoya, Sakao, Tamao, Takeuchi, Susumu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914337220722688
author Chijiwa, Daiki
Hasegawa, Taku
Nishida, Kyosuke
Yamaguchi, Shin'ya
Ohba, Tomoya
Sakao, Tamao
Takeuchi, Susumu
author_facet Chijiwa, Daiki
Hasegawa, Taku
Nishida, Kyosuke
Yamaguchi, Shin'ya
Ohba, Tomoya
Sakao, Tamao
Takeuchi, Susumu
contents Tokenization -- the process of decomposing a given text into a sequence of subwords called tokens -- is one of the key components in the development of language models. Particularly, auto-regressive language models generate texts token by token, i.e., by predicting the next-token distribution given the previous ones, and thus tokenization directly affects their efficiency in text generation. Since each language model has their own vocabulary as a set of possible tokens, they struggle to cooperate with each other at the level of next-token distributions such as model ensemble. In this paper, we establish a theoretical framework of lossless vocabulary reduction, which efficiently converts a given auto-regressive language model into the one with an arbitrarily small vocabulary without any loss in accuracy. This framework allows language models with different tokenization to cooperate with each other efficiently by reduction to their maximal common vocabulary. Specifically, we empirically demonstrate its applicability to model ensemble with different tokenization.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lossless Vocabulary Reduction for Auto-Regressive Language Models
Chijiwa, Daiki
Hasegawa, Taku
Nishida, Kyosuke
Yamaguchi, Shin'ya
Ohba, Tomoya
Sakao, Tamao
Takeuchi, Susumu
Computation and Language
Artificial Intelligence
Machine Learning
Tokenization -- the process of decomposing a given text into a sequence of subwords called tokens -- is one of the key components in the development of language models. Particularly, auto-regressive language models generate texts token by token, i.e., by predicting the next-token distribution given the previous ones, and thus tokenization directly affects their efficiency in text generation. Since each language model has their own vocabulary as a set of possible tokens, they struggle to cooperate with each other at the level of next-token distributions such as model ensemble. In this paper, we establish a theoretical framework of lossless vocabulary reduction, which efficiently converts a given auto-regressive language model into the one with an arbitrarily small vocabulary without any loss in accuracy. This framework allows language models with different tokenization to cooperate with each other efficiently by reduction to their maximal common vocabulary. Specifically, we empirically demonstrate its applicability to model ensemble with different tokenization.
title Lossless Vocabulary Reduction for Auto-Regressive Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.08102