Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909658158989312 |
|---|---|
| author | Karthika, N J Brahma, Maharaj Saluja, Rohit Ramakrishnan, Ganesh Desarkar, Maunendra Sankar |
| author_facet | Karthika, N J Brahma, Maharaj Saluja, Rohit Ramakrishnan, Ganesh Desarkar, Maunendra Sankar |
| contents | Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those in the Indian subcontinent. This paper presents a comprehensive intrinsic evaluation of tokenization strategies across 17 Indian languages. We quantify the trade-offs between bottom-up and top-down tokenizer algorithms (BPE and Unigram LM), effects of vocabulary sizes, and compare strategies of multilingual vocabulary construction such as joint and cluster-based training. We also show that extremely low-resource languages can benefit from tokenizers trained on related high-resource languages. Our study provides practical insights for building more fair, efficient, and linguistically informed tokenizers for multilingual NLP. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_17789 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights Karthika, N J Brahma, Maharaj Saluja, Rohit Ramakrishnan, Ganesh Desarkar, Maunendra Sankar Computation and Language Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those in the Indian subcontinent. This paper presents a comprehensive intrinsic evaluation of tokenization strategies across 17 Indian languages. We quantify the trade-offs between bottom-up and top-down tokenizer algorithms (BPE and Unigram LM), effects of vocabulary sizes, and compare strategies of multilingual vocabulary construction such as joint and cluster-based training. We also show that extremely low-resource languages can benefit from tokenizers trained on related high-resource languages. Our study provides practical insights for building more fair, efficient, and linguistically informed tokenizers for multilingual NLP. |
| title | Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2506.17789 |