Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Alakeel, Yara, Qwaider, Chatrine, Aldarmaki, Hanan, Alqahtani, Sawsan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
Automatic Restoration of Diacritics for Speech Data Sets
by: Shatnawi, Sara, et al.
Published: (2023)
by: Shatnawi, Sara, et al.
Published: (2023)
SparQLe: Speech Queries to Text Translation Through LLMs
by: Djanibekov, Amirbek, et al.
Published: (2025)
by: Djanibekov, Amirbek, et al.
Published: (2025)
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
by: Sayeed, Mohammad Amaan, et al.
Published: (2023)
by: Sayeed, Mohammad Amaan, et al.
Published: (2023)
ARWI: Arabic Write and Improve
by: Chirkunov, Kirill, et al.
Published: (2025)
by: Chirkunov, Kirill, et al.
Published: (2025)
Enhancing Arabic Automated Essay Scoring with Synthetic Data and Error Injection
by: Qwaider, Chatrine, et al.
Published: (2025)
by: Qwaider, Chatrine, et al.
Published: (2025)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
JEEM: Vision-Language Understanding in Four Arabic Dialects
by: Kadaoui, Karima, et al.
Published: (2025)
by: Kadaoui, Karima, et al.
Published: (2025)
Linear Semantic Segmentation for Low-Resource Spoken Dialects
by: Chirkunov, Kirill, et al.
Published: (2026)
by: Chirkunov, Kirill, et al.
Published: (2026)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
by: Alqahtani, Sawsan, et al.
Published: (2026)
by: Alqahtani, Sawsan, et al.
Published: (2026)
Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings
by: Alqahtani, Sawsan, et al.
Published: (2021)
by: Alqahtani, Sawsan, et al.
Published: (2021)
Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation
by: Wu, Sophie, et al.
Published: (2026)
by: Wu, Sophie, et al.
Published: (2026)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
by: Toyin, Hawau Olamide, et al.
Published: (2024)
by: Toyin, Hawau Olamide, et al.
Published: (2024)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
Personal Attribute Leakage in Federated Speech Models
by: Al-Ali, Hamdan, et al.
Published: (2025)
by: Al-Ali, Hamdan, et al.
Published: (2025)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026)
by: Nadeem, Afrozah, et al.
Published: (2026)
ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models
by: Alhumud, Anas, et al.
Published: (2026)
by: Alhumud, Anas, et al.
Published: (2026)
How well can LLMs Grade Essays in Arabic?
by: Ghazawi, Rayed, et al.
Published: (2025)
by: Ghazawi, Rayed, et al.
Published: (2025)
How Well Do LLMs Understand Tunisian Arabic?
by: Mahdi, Mohamed
Published: (2025)
by: Mahdi, Mohamed
Published: (2025)
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
by: Saadi, Nada, et al.
Published: (2025)
by: Saadi, Nada, et al.
Published: (2025)
GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning
by: Chouikhi, Hasna, et al.
Published: (2024)
by: Chouikhi, Hasna, et al.
Published: (2024)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
Enabling Scalable Evaluation of Bias Patterns in Medical LLMs
by: Fayyaz, Hamed, et al.
Published: (2024)
by: Fayyaz, Hamed, et al.
Published: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
by: Mustapha, Ahmad, et al.
Published: (2024)
by: Mustapha, Ahmad, et al.
Published: (2024)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
by: Baxi, Jatayu, et al.
Published: (2021)
by: Baxi, Jatayu, et al.
Published: (2021)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
by: Alshaikh, Rana, et al.
Published: (2025)
by: Alshaikh, Rana, et al.
Published: (2025)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Clinical Annotations for Automatic Stuttering Severity Assessment
by: Valente, Ana Rita, et al.
Published: (2025)
by: Valente, Ana Rita, et al.
Published: (2025)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
by: Chen, Haoyang, et al.
Published: (2026)
by: Chen, Haoyang, et al.
Published: (2026)
Towards stable AI systems for Evaluating Arabic Pronunciations
by: Zaatiti, Hadi, et al.
Published: (2025)
by: Zaatiti, Hadi, et al.
Published: (2025)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
by: Sarangi, Sneheel, et al.
Published: (2025)
by: Sarangi, Sneheel, et al.
Published: (2025)
Slimming Down LLMs Without Losing Their Minds
by: Qingda, et al.
Published: (2025)
by: Qingda, et al.
Published: (2025)
GRDD: A Dataset for Greek Dialectal NLP
by: Chatzikyriakidis, Stergios, et al.
Published: (2023)
by: Chatzikyriakidis, Stergios, et al.
Published: (2023)
AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
by: Kashirskiy, Mark, et al.
Published: (2025)
by: Kashirskiy, Mark, et al.
Published: (2025)
LLMs are Not Just Next Token Predictors
by: Downes, Stephen M., et al.
Published: (2024)
by: Downes, Stephen M., et al.
Published: (2024)
The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness
by: Shaban, Sanad, et al.
Published: (2025)
by: Shaban, Sanad, et al.
Published: (2025)
SalamahBench: Toward Standardized Safety Evaluation for Arabic Language Models
by: Abdelnasser, Omar, et al.
Published: (2026)
by: Abdelnasser, Omar, et al.
Published: (2026)
Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation
by: Bellaouar, Slimane, et al.
Published: (2025)
by: Bellaouar, Slimane, et al.
Published: (2025)
C3AI: Crafting and Evaluating Constitutions for Constitutional AI
by: Kyrychenko, Yara, et al.
Published: (2025)
by: Kyrychenko, Yara, et al.
Published: (2025)
Similar Items
-
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
by: Toyin, Hawau Olamide, et al.
Published: (2025) -
Automatic Restoration of Diacritics for Speech Data Sets
by: Shatnawi, Sara, et al.
Published: (2023) -
SparQLe: Speech Queries to Text Translation Through LLMs
by: Djanibekov, Amirbek, et al.
Published: (2025) -
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
by: Sayeed, Mohammad Amaan, et al.
Published: (2023) -
ARWI: Arabic Write and Improve
by: Chirkunov, Kirill, et al.
Published: (2025)