Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chatzi, Ivi, Benz, Nina Corvelo, Tsirtsis, Stratis, Gomez-Rodriguez, Manuel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918315049353216
author Chatzi, Ivi
Benz, Nina Corvelo
Tsirtsis, Stratis
Gomez-Rodriguez, Manuel
author_facet Chatzi, Ivi
Benz, Nina Corvelo
Tsirtsis, Stratis
Gomez-Rodriguez, Manuel
contents Providers of LLM-as-a-service have predominantly adopted a simple pricing model: users pay a fixed price per token. Consequently, one may think that the price two different users would pay for the same output string under the same input prompt is the same. In our work, we show that, surprisingly, this is not (always) true. We find empirical evidence that, particularly for non-english outputs, both proprietary and open-weights LLMs often generate the same (output) string with multiple different tokenizations, even under the same input prompt, and this in turn leads to arbitrary price variation. To address the problem of tokenization multiplicity, we introduce canonical generation, a type of constrained generation that restricts LLMs to only generate canonical tokenizations -- the unique tokenization in which each string is tokenized during the training process of an LLM. Further, we introduce an efficient sampling algorithm for canonical generation based on the Gumbel-Max trick. Experiments on a variety of natural language tasks demonstrate that our sampling algorithm for canonical generation is comparable to standard sampling in terms of performance and runtime, and it solves the problem of tokenization multiplicity.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
Chatzi, Ivi
Benz, Nina Corvelo
Tsirtsis, Stratis
Gomez-Rodriguez, Manuel
Computation and Language
Artificial Intelligence
Machine Learning
Providers of LLM-as-a-service have predominantly adopted a simple pricing model: users pay a fixed price per token. Consequently, one may think that the price two different users would pay for the same output string under the same input prompt is the same. In our work, we show that, surprisingly, this is not (always) true. We find empirical evidence that, particularly for non-english outputs, both proprietary and open-weights LLMs often generate the same (output) string with multiple different tokenizations, even under the same input prompt, and this in turn leads to arbitrary price variation. To address the problem of tokenization multiplicity, we introduce canonical generation, a type of constrained generation that restricts LLMs to only generate canonical tokenizations -- the unique tokenization in which each string is tokenized during the training process of an LLM. Further, we introduce an efficient sampling algorithm for canonical generation based on the Gumbel-Max trick. Experiments on a variety of natural language tasks demonstrate that our sampling algorithm for canonical generation is comparable to standard sampling in terms of performance and runtime, and it solves the problem of tokenization multiplicity.
title Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.06446