Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karthika, N J, Brahma, Maharaj, Saluja, Rohit, Ramakrishnan, Ganesh, Desarkar, Maunendra Sankar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909658158989312
author Karthika, N J
Brahma, Maharaj
Saluja, Rohit
Ramakrishnan, Ganesh
Desarkar, Maunendra Sankar
author_facet Karthika, N J
Brahma, Maharaj
Saluja, Rohit
Ramakrishnan, Ganesh
Desarkar, Maunendra Sankar
contents Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those in the Indian subcontinent. This paper presents a comprehensive intrinsic evaluation of tokenization strategies across 17 Indian languages. We quantify the trade-offs between bottom-up and top-down tokenizer algorithms (BPE and Unigram LM), effects of vocabulary sizes, and compare strategies of multilingual vocabulary construction such as joint and cluster-based training. We also show that extremely low-resource languages can benefit from tokenizers trained on related high-resource languages. Our study provides practical insights for building more fair, efficient, and linguistically informed tokenizers for multilingual NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
Karthika, N J
Brahma, Maharaj
Saluja, Rohit
Ramakrishnan, Ganesh
Desarkar, Maunendra Sankar
Computation and Language
Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those in the Indian subcontinent. This paper presents a comprehensive intrinsic evaluation of tokenization strategies across 17 Indian languages. We quantify the trade-offs between bottom-up and top-down tokenizer algorithms (BPE and Unigram LM), effects of vocabulary sizes, and compare strategies of multilingual vocabulary construction such as joint and cluster-based training. We also show that extremely low-resource languages can benefit from tokenizers trained on related high-resource languages. Our study provides practical insights for building more fair, efficient, and linguistically informed tokenizers for multilingual NLP.
title Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
topic Computation and Language
url https://arxiv.org/abs/2506.17789