Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vemula, Saketh Reddy, Dandapat, Sandipan, Sharma, Dipti Misra, Krishnamurthy, Parameswari |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
von: Asgari, Ehsaneddin, et al.
Veröffentlicht: (2025)
von: Asgari, Ehsaneddin, et al.
Veröffentlicht: (2025)
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024)
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024)
Batching BPE Tokenization Merges
von: Morgan, Alexander P.
Veröffentlicht: (2024)
von: Morgan, Alexander P.
Veröffentlicht: (2024)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)
H-Net++: Hierarchical Dynamic Chunking for Tokenizer-Free Language Modelling in Morphologically-Rich Languages
von: Zakershahrak, Mehrdad, et al.
Veröffentlicht: (2025)
von: Zakershahrak, Mehrdad, et al.
Veröffentlicht: (2025)
The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles
von: M, Abhinav P, et al.
Veröffentlicht: (2025)
von: M, Abhinav P, et al.
Veröffentlicht: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
von: Bala, Abhinaba, et al.
Veröffentlicht: (2024)
von: Bala, Abhinaba, et al.
Veröffentlicht: (2024)
Fine-tuning Pre-trained Named Entity Recognition Models For Indian Languages
von: Bahad, Sankalp, et al.
Veröffentlicht: (2024)
von: Bahad, Sankalp, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
von: Kadamba, Venu Gopal, et al.
Veröffentlicht: (2026)
von: Kadamba, Venu Gopal, et al.
Veröffentlicht: (2026)
Decoding Fake Narratives in Spreading Hateful Stories: A Dual-Head RoBERTa Model with Multi-Task Learning
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
von: Alakeel, Yara, et al.
Veröffentlicht: (2026)
von: Alakeel, Yara, et al.
Veröffentlicht: (2026)
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2023)
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2023)
Conditional Unigram Tokenization with Parallel Data
von: Vico, Gianluca, et al.
Veröffentlicht: (2025)
von: Vico, Gianluca, et al.
Veröffentlicht: (2025)
LLM Safety for Children
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
von: Rath, Prasanjit, et al.
Veröffentlicht: (2025)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
von: Parra, Iñigo
Veröffentlicht: (2024)
von: Parra, Iñigo
Veröffentlicht: (2024)
Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
von: Xia, Chengxuan, et al.
Veröffentlicht: (2025)
von: Xia, Chengxuan, et al.
Veröffentlicht: (2025)
QuranMorph: Morphologically Annotated Quranic Corpus
von: Akra, Diyam, et al.
Veröffentlicht: (2025)
von: Akra, Diyam, et al.
Veröffentlicht: (2025)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
von: Thakur, Aamod, et al.
Veröffentlicht: (2025)
von: Thakur, Aamod, et al.
Veröffentlicht: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
BlockBPE: Parallel BPE Tokenization
von: You, Amos
Veröffentlicht: (2025)
von: You, Amos
Veröffentlicht: (2025)
State over Tokens: Characterizing the Role of Reasoning Tokens
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
von: Levy, Mosh, et al.
Veröffentlicht: (2025)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
von: Krishnamurthy, Vikram
Veröffentlicht: (2026)
von: Krishnamurthy, Vikram
Veröffentlicht: (2026)
Morphological evaluation of subwords vocabulary used by BETO language model
von: García-Sierra, Óscar, et al.
Veröffentlicht: (2024)
von: García-Sierra, Óscar, et al.
Veröffentlicht: (2024)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
von: Akula, Arjun R., et al.
Veröffentlicht: (2025)
von: Akula, Arjun R., et al.
Veröffentlicht: (2025)
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation
von: Sälevä, Jonne, et al.
Veröffentlicht: (2021)
von: Sälevä, Jonne, et al.
Veröffentlicht: (2021)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Token Alignment via Character Matching for Subword Completion
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies
von: Sun, Xin, et al.
Veröffentlicht: (2024)
von: Sun, Xin, et al.
Veröffentlicht: (2024)
Rethinking Adapter Placement: A Dominant Adaptation Module Perspective
von: Zhang, Suoxin, et al.
Veröffentlicht: (2026)
von: Zhang, Suoxin, et al.
Veröffentlicht: (2026)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
von: Bhattacharjee, Soham, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Soham, et al.
Veröffentlicht: (2025)
Shona spaCy: A Morphological Analyzer for an Under-Resourced Bantu Language
von: Masoka, Happymore
Veröffentlicht: (2025)
von: Masoka, Happymore
Veröffentlicht: (2025)
ERAS: Evaluating the Robustness of Chinese NLP Models to Morphological Garden Path Errors
von: Li, Qinchan, et al.
Veröffentlicht: (2024)
von: Li, Qinchan, et al.
Veröffentlicht: (2024)
From Language Models over Tokens to Language Models over Characters
von: Vieira, Tim, et al.
Veröffentlicht: (2024)
von: Vieira, Tim, et al.
Veröffentlicht: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
von: Shi, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Shi, Xiaofeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025) -
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
von: Asgari, Ehsaneddin, et al.
Veröffentlicht: (2025) -
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024) -
Batching BPE Tokenization Merges
von: Morgan, Alexander P.
Veröffentlicht: (2024) -
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
von: Bhaskar, Yash, et al.
Veröffentlicht: (2025)