Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Vemula, Saketh Reddy, Dandapat, Sandipan, Sharma, Dipti Misra, Krishnamurthy, Parameswari |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025)
by: Asgari, Ehsaneddin, et al.
Published: (2025)
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
by: Mujadia, Vandan, et al.
Published: (2024)
by: Mujadia, Vandan, et al.
Published: (2024)
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024)
by: Morgan, Alexander P.
Published: (2024)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
H-Net++: Hierarchical Dynamic Chunking for Tokenizer-Free Language Modelling in Morphologically-Rich Languages
by: Zakershahrak, Mehrdad, et al.
Published: (2025)
by: Zakershahrak, Mehrdad, et al.
Published: (2025)
The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles
by: M, Abhinav P, et al.
Published: (2025)
by: M, Abhinav P, et al.
Published: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
by: Bala, Abhinaba, et al.
Published: (2024)
by: Bala, Abhinaba, et al.
Published: (2024)
Fine-tuning Pre-trained Named Entity Recognition Models For Indian Languages
by: Bahad, Sankalp, et al.
Published: (2024)
by: Bahad, Sankalp, et al.
Published: (2024)
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
by: Alakeel, Yara, et al.
Published: (2026)
by: Alakeel, Yara, et al.
Published: (2026)
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages
by: Niyogi, Mitodru, et al.
Published: (2024)
by: Niyogi, Mitodru, et al.
Published: (2024)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
by: Vashishtha, Aniket, et al.
Published: (2023)
by: Vashishtha, Aniket, et al.
Published: (2023)
Conditional Unigram Tokenization with Parallel Data
by: Vico, Gianluca, et al.
Published: (2025)
by: Vico, Gianluca, et al.
Published: (2025)
LLM Safety for Children
by: Rath, Prasanjit, et al.
Published: (2025)
by: Rath, Prasanjit, et al.
Published: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
by: Xia, Chengxuan, et al.
Published: (2025)
by: Xia, Chengxuan, et al.
Published: (2025)
QuranMorph: Morphologically Annotated Quranic Corpus
by: Akra, Diyam, et al.
Published: (2025)
by: Akra, Diyam, et al.
Published: (2025)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
by: Thakur, Aamod, et al.
Published: (2025)
by: Thakur, Aamod, et al.
Published: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025)
by: You, Amos
Published: (2025)
State over Tokens: Characterizing the Role of Reasoning Tokens
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
by: Krishnamurthy, Vikram
Published: (2026)
by: Krishnamurthy, Vikram
Published: (2026)
Morphological evaluation of subwords vocabulary used by BETO language model
by: García-Sierra, Óscar, et al.
Published: (2024)
by: García-Sierra, Óscar, et al.
Published: (2024)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
by: Akula, Arjun R., et al.
Published: (2025)
by: Akula, Arjun R., et al.
Published: (2025)
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation
by: Sälevä, Jonne, et al.
Published: (2021)
by: Sälevä, Jonne, et al.
Published: (2021)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
by: Hu, Shijing, et al.
Published: (2025)
by: Hu, Shijing, et al.
Published: (2025)
Token Alignment via Character Matching for Subword Completion
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies
by: Sun, Xin, et al.
Published: (2024)
by: Sun, Xin, et al.
Published: (2024)
Rethinking Adapter Placement: A Dominant Adaptation Module Perspective
by: Zhang, Suoxin, et al.
Published: (2026)
by: Zhang, Suoxin, et al.
Published: (2026)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
by: Bhattacharjee, Soham, et al.
Published: (2025)
by: Bhattacharjee, Soham, et al.
Published: (2025)
Shona spaCy: A Morphological Analyzer for an Under-Resourced Bantu Language
by: Masoka, Happymore
Published: (2025)
by: Masoka, Happymore
Published: (2025)
ERAS: Evaluating the Robustness of Chinese NLP Models to Morphological Garden Path Errors
by: Li, Qinchan, et al.
Published: (2024)
by: Li, Qinchan, et al.
Published: (2024)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
by: Shi, Xiaofeng, et al.
Published: (2025)
by: Shi, Xiaofeng, et al.
Published: (2025)
Similar Items
-
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
by: Vemula, Saketh Reddy, et al.
Published: (2025) -
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025) -
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
by: Mujadia, Vandan, et al.
Published: (2024) -
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024) -
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
by: Bhaskar, Yash, et al.
Published: (2025)