MUTANT: A Recipe for Multilingual Tokenizer Design

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rana, Souvik, Menezes, Arul, Kulkarni, Ashish, Khatri, Chandra, Agarwal, Shubham
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914412615434240
author Rana, Souvik
Menezes, Arul
Kulkarni, Ashish
Khatri, Chandra
Agarwal, Shubham
author_facet Rana, Souvik
Menezes, Arul
Kulkarni, Ashish
Khatri, Chandra
Agarwal, Shubham
contents Tokenizers play a crucial role in determining the performance, training efficiency, and the inference cost of Large Language Models (LLMs). Designing effective tokenizers for multilingual LLMs is particularly challenging due to diverse scripts and rich morphological variation. While subword methods like Byte Pair Encoding (BPE) are widely adopted, their effectiveness in multilingual settings remains underexplored. We present MUTANT, a recipe for building multilingual tokenizers, with careful vocabulary and training data design, language-aware pre-tokenization, and subword and multiword aware training. We also introduce MUTANT-Indic, a tokenizer for India-specific multilingual LLMs, that produces linguistically coherent tokens and achieves state-of-the-art performance. Evaluated across English, 22 Indian languages and code data, our tokenizer improves the average fertility score by 39.5%$ over LLaMA4 and by 18% over Sutra (the current best). This translates to 44% improvement in inference throughput over LLaMA4 while maintaining comparable performance on English and Indic benchmarks. We present detailed ablations across tokenizer training data size, vocabulary size, merging techniques, and pre-tokenization strategies, demonstrating the robustness of our design choices.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03237
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MUTANT: A Recipe for Multilingual Tokenizer Design
Rana, Souvik
Menezes, Arul
Kulkarni, Ashish
Khatri, Chandra
Agarwal, Shubham
Computation and Language
Tokenizers play a crucial role in determining the performance, training efficiency, and the inference cost of Large Language Models (LLMs). Designing effective tokenizers for multilingual LLMs is particularly challenging due to diverse scripts and rich morphological variation. While subword methods like Byte Pair Encoding (BPE) are widely adopted, their effectiveness in multilingual settings remains underexplored. We present MUTANT, a recipe for building multilingual tokenizers, with careful vocabulary and training data design, language-aware pre-tokenization, and subword and multiword aware training. We also introduce MUTANT-Indic, a tokenizer for India-specific multilingual LLMs, that produces linguistically coherent tokens and achieves state-of-the-art performance. Evaluated across English, 22 Indian languages and code data, our tokenizer improves the average fertility score by 39.5%$ over LLaMA4 and by 18% over Sutra (the current best). This translates to 44% improvement in inference throughput over LLaMA4 while maintaining comparable performance on English and Indic benchmarks. We present detailed ablations across tokenizer training data size, vocabulary size, merging techniques, and pre-tokenization strategies, demonstrating the robustness of our design choices.
title MUTANT: A Recipe for Multilingual Tokenizer Design
topic Computation and Language
url https://arxiv.org/abs/2511.03237