MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Asgari, Ehsaneddin, Kheir, Yassine El, Javaheri, Mohammad Ali Sadraei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SuperPos-Prompt: Enhancing Soft Prompt Tuning of Language Models with Superposition of Multi Token Embeddings
von: SadraeiJavaeri, MohammadAli, et al.
Veröffentlicht: (2024)
von: SadraeiJavaeri, MohammadAli, et al.
Veröffentlicht: (2024)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
Batching BPE Tokenization Merges
von: Morgan, Alexander P.
Veröffentlicht: (2024)
von: Morgan, Alexander P.
Veröffentlicht: (2024)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2024)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2024)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)
TuringQ: Benchmarking AI Comprehension in Theory of Computation
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024)
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024)
EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
von: Nazar, Nizi, et al.
Veröffentlicht: (2025)
von: Nazar, Nizi, et al.
Veröffentlicht: (2025)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2025)
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2025)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024)
von: Chizhov, Pavel, et al.
Veröffentlicht: (2024)
BlockBPE: Parallel BPE Tokenization
von: You, Amos
Veröffentlicht: (2025)
von: You, Amos
Veröffentlicht: (2025)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
von: Mofayezi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Mofayezi, Mohammadreza, et al.
Veröffentlicht: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
von: Kadamba, Venu Gopal, et al.
Veröffentlicht: (2026)
von: Kadamba, Venu Gopal, et al.
Veröffentlicht: (2026)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2026)
AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis
von: Kure, Alireza Ghahramani, et al.
Veröffentlicht: (2025)
von: Kure, Alireza Ghahramani, et al.
Veröffentlicht: (2025)
AIMA at SemEval-2024 Task 10: History-Based Emotion Recognition in Hindi-English Code-Mixed Conversations
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
MorphNAS: Differentiable Architecture Search for Morphologically-Aware Multilingual NER
von: Devadiga, Prathamesh, et al.
Veröffentlicht: (2025)
von: Devadiga, Prathamesh, et al.
Veröffentlicht: (2025)
QuranMorph: Morphologically Annotated Quranic Corpus
von: Akra, Diyam, et al.
Veröffentlicht: (2025)
von: Akra, Diyam, et al.
Veröffentlicht: (2025)
Constructing a BPE Tokenization DFA
von: Berglund, Martin, et al.
Veröffentlicht: (2024)
von: Berglund, Martin, et al.
Veröffentlicht: (2024)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
von: Zheng, Weihua, et al.
Veröffentlicht: (2026)
von: Zheng, Weihua, et al.
Veröffentlicht: (2026)
AlphaToken: Decoupling Adaptation and Stability for Path-Aware Response Token Valuation in LLM Post-Training
von: Qing, Liu, et al.
Veröffentlicht: (2026)
von: Qing, Liu, et al.
Veröffentlicht: (2026)
Linguistics-Aware Non-Distortionary LLM Watermarking
von: Park, Shinwoo, et al.
Veröffentlicht: (2026)
von: Park, Shinwoo, et al.
Veröffentlicht: (2026)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
von: Daouadi, Kheir Eddine, et al.
Veröffentlicht: (2024)
von: Daouadi, Kheir Eddine, et al.
Veröffentlicht: (2024)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Personality Expression Across Contexts: Linguistic and Behavioral Variation in LLM Agents
von: Han, Bin, et al.
Veröffentlicht: (2026)
von: Han, Bin, et al.
Veröffentlicht: (2026)
MorphPiece : A Linguistic Tokenizer for Large Language Models
von: Jabbar, Haris
Veröffentlicht: (2023)
von: Jabbar, Haris
Veröffentlicht: (2023)
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
von: Cekinmez, Jasin, et al.
Veröffentlicht: (2025)
von: Cekinmez, Jasin, et al.
Veröffentlicht: (2025)
Byte BPE Tokenization as an Inverse string Homomorphism
von: Geng, Saibo, et al.
Veröffentlicht: (2024)
von: Geng, Saibo, et al.
Veröffentlicht: (2024)
MorphTok: Morphologically Grounded Tokenization for Indian Languages
von: Brahma, Maharaj, et al.
Veröffentlicht: (2025)
von: Brahma, Maharaj, et al.
Veröffentlicht: (2025)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
von: Park, Shinwoo, et al.
Veröffentlicht: (2025)
von: Park, Shinwoo, et al.
Veröffentlicht: (2025)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
von: Sun, Yike, et al.
Veröffentlicht: (2026)
von: Sun, Yike, et al.
Veröffentlicht: (2026)
Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries
von: Zhang, Yuchen, et al.
Veröffentlicht: (2026)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2026)
Morphological Typology in BPE Subword Productivity and Language Modeling
von: Parra, Iñigo
Veröffentlicht: (2024)
von: Parra, Iñigo
Veröffentlicht: (2024)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
Training-Trajectory-Aware Token Selection
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
von: Dolga, Rares, et al.
Veröffentlicht: (2025)
von: Dolga, Rares, et al.
Veröffentlicht: (2025)
AdaptBPE: From General Purpose to Specialized Tokenizers
von: Liyanage, Vijini, et al.
Veröffentlicht: (2026)
von: Liyanage, Vijini, et al.
Veröffentlicht: (2026)
Training Text-to-Molecule Models with Context-Aware Tokenization
von: Kim, Seojin, et al.
Veröffentlicht: (2025)
von: Kim, Seojin, et al.
Veröffentlicht: (2025)
Token-Budget-Aware LLM Reasoning
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SuperPos-Prompt: Enhancing Soft Prompt Tuning of Language Models with Superposition of Multi Token Embeddings
von: SadraeiJavaeri, MohammadAli, et al.
Veröffentlicht: (2024) -
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025) -
Batching BPE Tokenization Merges
von: Morgan, Alexander P.
Veröffentlicht: (2024) -
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
von: Abootorabi, Mohammad Mahdi, et al.
Veröffentlicht: (2024) -
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
von: Ghahroodi, Omid, et al.
Veröffentlicht: (2024)