Is There a Case for Conversation Optimized Tokenizers in Large Language Models?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ferrando, Raquel, Conde, Javier, Martínez, Gonzalo, Reviriego, Pedro
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911019415109632
author Ferrando, Raquel
Conde, Javier
Martínez, Gonzalo
Reviriego, Pedro
author_facet Ferrando, Raquel
Conde, Javier
Martínez, Gonzalo
Reviriego, Pedro
contents The computational and energy costs of Large Language Models (LLMs) have increased exponentially driven by the growing model sizes and the massive adoption of LLMs by hundreds of millions of users. The unit cost of an LLM is the computation of a token. Therefore, the tokenizer plays an important role in the efficiency of a model, and they are carefully optimized to minimize the number of tokens for the text in their training corpus. One of the most popular applications of LLMs are chatbots that interact with users. A key observation is that, for those chatbots, what is important is the performance of the tokenizer in the user text input and the chatbot responses. Those are most likely different from the text in the training corpus. So, a question that immediately arises is whether there is a potential benefit in optimizing tokenizers for chatbot conversations. In this paper, this idea is explored for different tokenizers by using a publicly available corpus of chatbot conversations to redesign their vocabularies and evaluate their performance in this domain. The results show that conversation-optimized tokenizers consistently reduce the number of tokens in chatbot dialogues, which can lead to meaningful energy savings, in the range of 5% to 10% while having minimal or even slightly positive impact on tokenization efficiency for the original training corpus.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
Ferrando, Raquel
Conde, Javier
Martínez, Gonzalo
Reviriego, Pedro
Computation and Language
Artificial Intelligence
The computational and energy costs of Large Language Models (LLMs) have increased exponentially driven by the growing model sizes and the massive adoption of LLMs by hundreds of millions of users. The unit cost of an LLM is the computation of a token. Therefore, the tokenizer plays an important role in the efficiency of a model, and they are carefully optimized to minimize the number of tokens for the text in their training corpus. One of the most popular applications of LLMs are chatbots that interact with users. A key observation is that, for those chatbots, what is important is the performance of the tokenizer in the user text input and the chatbot responses. Those are most likely different from the text in the training corpus. So, a question that immediately arises is whether there is a potential benefit in optimizing tokenizers for chatbot conversations. In this paper, this idea is explored for different tokenizers by using a publicly available corpus of chatbot conversations to redesign their vocabularies and evaluate their performance in this domain. The results show that conversation-optimized tokenizers consistently reduce the number of tokens in chatbot dialogues, which can lead to meaningful energy savings, in the range of 5% to 10% while having minimal or even slightly positive impact on tokenization efficiency for the original training corpus.
title Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.18674