Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Vemula, Saketh Reddy, Dandapat, Sandipan, Sharma, Dipti Misra, Krishnamurthy, Parameswari |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
par: Vemula, Saketh Reddy, et autres
Publié: (2025)
par: Vemula, Saketh Reddy, et autres
Publié: (2025)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
par: Asgari, Ehsaneddin, et autres
Publié: (2025)
par: Asgari, Ehsaneddin, et autres
Publié: (2025)
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
par: Mujadia, Vandan, et autres
Publié: (2024)
par: Mujadia, Vandan, et autres
Publié: (2024)
Batching BPE Tokenization Merges
par: Morgan, Alexander P.
Publié: (2024)
par: Morgan, Alexander P.
Publié: (2024)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
par: Bhaskar, Yash, et autres
Publié: (2025)
par: Bhaskar, Yash, et autres
Publié: (2025)
H-Net++: Hierarchical Dynamic Chunking for Tokenizer-Free Language Modelling in Morphologically-Rich Languages
par: Zakershahrak, Mehrdad, et autres
Publié: (2025)
par: Zakershahrak, Mehrdad, et autres
Publié: (2025)
The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles
par: M, Abhinav P, et autres
Publié: (2025)
par: M, Abhinav P, et autres
Publié: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
par: Mittal, Avni, et autres
Publié: (2026)
par: Mittal, Avni, et autres
Publié: (2026)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
par: Kumar, Shanu, et autres
Publié: (2024)
par: Kumar, Shanu, et autres
Publié: (2024)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
par: Bala, Abhinaba, et autres
Publié: (2024)
par: Bala, Abhinaba, et autres
Publié: (2024)
Fine-tuning Pre-trained Named Entity Recognition Models For Indian Languages
par: Bahad, Sankalp, et autres
Publié: (2024)
par: Bahad, Sankalp, et autres
Publié: (2024)
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
par: Kumar, Shanu, et autres
Publié: (2024)
par: Kumar, Shanu, et autres
Publié: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
par: Kadamba, Venu Gopal, et autres
Publié: (2026)
par: Kadamba, Venu Gopal, et autres
Publié: (2026)
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
par: Bhaskar, Yash, et autres
Publié: (2025)
par: Bhaskar, Yash, et autres
Publié: (2025)
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
par: Alakeel, Yara, et autres
Publié: (2026)
par: Alakeel, Yara, et autres
Publié: (2026)
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages
par: Niyogi, Mitodru, et autres
Publié: (2024)
par: Niyogi, Mitodru, et autres
Publié: (2024)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
par: Vashishtha, Aniket, et autres
Publié: (2023)
par: Vashishtha, Aniket, et autres
Publié: (2023)
Conditional Unigram Tokenization with Parallel Data
par: Vico, Gianluca, et autres
Publié: (2025)
par: Vico, Gianluca, et autres
Publié: (2025)
LLM Safety for Children
par: Rath, Prasanjit, et autres
Publié: (2025)
par: Rath, Prasanjit, et autres
Publié: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
par: Shrestha, Adarsha, et autres
Publié: (2025)
par: Shrestha, Adarsha, et autres
Publié: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
par: Parra, Iñigo
Publié: (2024)
par: Parra, Iñigo
Publié: (2024)
Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
par: Xia, Chengxuan, et autres
Publié: (2025)
par: Xia, Chengxuan, et autres
Publié: (2025)
QuranMorph: Morphologically Annotated Quranic Corpus
par: Akra, Diyam, et autres
Publié: (2025)
par: Akra, Diyam, et autres
Publié: (2025)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
par: Thakur, Aamod, et autres
Publié: (2025)
par: Thakur, Aamod, et autres
Publié: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
par: Ismayilzada, Mete, et autres
Publié: (2024)
par: Ismayilzada, Mete, et autres
Publié: (2024)
BlockBPE: Parallel BPE Tokenization
par: You, Amos
Publié: (2025)
par: You, Amos
Publié: (2025)
State over Tokens: Characterizing the Role of Reasoning Tokens
par: Levy, Mosh, et autres
Publié: (2025)
par: Levy, Mosh, et autres
Publié: (2025)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
par: Krishnamurthy, Vikram
Publié: (2026)
par: Krishnamurthy, Vikram
Publié: (2026)
Morphological evaluation of subwords vocabulary used by BETO language model
par: García-Sierra, Óscar, et autres
Publié: (2024)
par: García-Sierra, Óscar, et autres
Publié: (2024)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
par: Akula, Arjun R., et autres
Publié: (2025)
par: Akula, Arjun R., et autres
Publié: (2025)
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation
par: Sälevä, Jonne, et autres
Publié: (2021)
par: Sälevä, Jonne, et autres
Publié: (2021)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
par: Hu, Shijing, et autres
Publié: (2025)
par: Hu, Shijing, et autres
Publié: (2025)
Token Alignment via Character Matching for Subword Completion
par: Athiwaratkun, Ben, et autres
Publié: (2024)
par: Athiwaratkun, Ben, et autres
Publié: (2024)
Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies
par: Sun, Xin, et autres
Publié: (2024)
par: Sun, Xin, et autres
Publié: (2024)
Rethinking Adapter Placement: A Dominant Adaptation Module Perspective
par: Zhang, Suoxin, et autres
Publié: (2026)
par: Zhang, Suoxin, et autres
Publié: (2026)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
par: Bhattacharjee, Soham, et autres
Publié: (2025)
par: Bhattacharjee, Soham, et autres
Publié: (2025)
Shona spaCy: A Morphological Analyzer for an Under-Resourced Bantu Language
par: Masoka, Happymore
Publié: (2025)
par: Masoka, Happymore
Publié: (2025)
ERAS: Evaluating the Robustness of Chinese NLP Models to Morphological Garden Path Errors
par: Li, Qinchan, et autres
Publié: (2024)
par: Li, Qinchan, et autres
Publié: (2024)
From Language Models over Tokens to Language Models over Characters
par: Vieira, Tim, et autres
Publié: (2024)
par: Vieira, Tim, et autres
Publié: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
par: Shi, Xiaofeng, et autres
Publié: (2025)
par: Shi, Xiaofeng, et autres
Publié: (2025)
Documents similaires
-
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
par: Vemula, Saketh Reddy, et autres
Publié: (2025) -
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
par: Asgari, Ehsaneddin, et autres
Publié: (2025) -
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
par: Mujadia, Vandan, et autres
Publié: (2024) -
Batching BPE Tokenization Merges
par: Morgan, Alexander P.
Publié: (2024) -
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
par: Bhaskar, Yash, et autres
Publié: (2025)