Separate Before You Compress: The WWHO Tokenization Architecture
Fuente:
arXiv
Saved in:
| Main Author: | Darshana, Kusal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025)
by: Zhang, Junwen, et al.
Published: (2025)
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Approaching I/O-optimality for Approximate Attention
by: Papp, Pál András, et al.
Published: (2026)
by: Papp, Pál András, et al.
Published: (2026)
Attention Meets Reachability: Structural Equivalence and Efficiency in Grammar-Constrained LLM Decoding
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
A First Runtime Analysis of the PAES-25: An Enhanced Variant of the Pareto Archived Evolution Strategy
by: Opris, Andre
Published: (2025)
by: Opris, Andre
Published: (2025)
Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs
by: Brown, Nik Bear
Published: (2024)
by: Brown, Nik Bear
Published: (2024)
How to Compute a Moving Sum
by: Maslen, David K., et al.
Published: (2025)
by: Maslen, David K., et al.
Published: (2025)
Latent Objective Induction and Diversity-Constrained Selection: Algorithms for Multi-Locale Retrieval Pipelines
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
by: Hasan, Md Arid, et al.
Published: (2025)
by: Hasan, Md Arid, et al.
Published: (2025)
PropXplain: Can LLMs Enable Explainable Propaganda Detection?
by: Hasanain, Maram, et al.
Published: (2025)
by: Hasanain, Maram, et al.
Published: (2025)
Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
by: De Mel, Yomal, et al.
Published: (2024)
by: De Mel, Yomal, et al.
Published: (2024)
Large Language Models for Propaganda Span Annotation
by: Hasanain, Maram, et al.
Published: (2023)
by: Hasanain, Maram, et al.
Published: (2023)
Tokenizations for Austronesian Language Models: study on languages in Indonesia Archipelago
by: Lumbantobing, Andhika Bernard, et al.
Published: (2026)
by: Lumbantobing, Andhika Bernard, et al.
Published: (2026)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
by: Lei, Xiang, et al.
Published: (2025)
by: Lei, Xiang, et al.
Published: (2025)
The Optimizer Quotient and the Certification Trilemma
by: Simas, Tristan
Published: (2026)
by: Simas, Tristan
Published: (2026)
Towards a Rigorous Understanding of the Population Dynamics of the NSGA-III: Tight Runtime Bounds
by: Opris, Andre
Published: (2025)
by: Opris, Andre
Published: (2025)
Stealth edits to large language models
by: Sutton, Oliver J., et al.
Published: (2024)
by: Sutton, Oliver J., et al.
Published: (2024)
Runtime Analyses of NSGA-III on Many-Objective Problems
by: Opris, Andre, et al.
Published: (2024)
by: Opris, Andre, et al.
Published: (2024)
Achieving Tight $O(4^k)$ Runtime Bounds on Jump$_k$ by Proving that Genetic Algorithms Evolve Near-Maximal Population Diversity
by: Opris, Andre, et al.
Published: (2024)
by: Opris, Andre, et al.
Published: (2024)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
by: Khan, Ishraq, et al.
Published: (2025)
by: Khan, Ishraq, et al.
Published: (2025)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
by: Guan, Xin, et al.
Published: (2024)
by: Guan, Xin, et al.
Published: (2024)
CELI: Controller-Embedded Language Model Interactions
by: Wagner, Jan-Samuel, et al.
Published: (2024)
by: Wagner, Jan-Samuel, et al.
Published: (2024)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026)
by: Haque, Md. Asraful, et al.
Published: (2026)
Optimizing Genetic Algorithms Using the Binomial Distribution
by: Cicirello, Vincent A.
Published: (2024)
by: Cicirello, Vincent A.
Published: (2024)
Uncovering Uncertainty in Transformer Inference
by: Brothers, Greyson, et al.
Published: (2024)
by: Brothers, Greyson, et al.
Published: (2024)
Multimodal Structure-Aware Quantum Data Processing
by: Hawashin, Hala, et al.
Published: (2024)
by: Hawashin, Hala, et al.
Published: (2024)
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking
by: Dalvi, Fahim, et al.
Published: (2023)
by: Dalvi, Fahim, et al.
Published: (2023)
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
by: Mousi, Basel, et al.
Published: (2024)
by: Mousi, Basel, et al.
Published: (2024)
NativQA Framework: Enabling LLMs and VLMs with Native, Local, and Everyday Knowledge
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
by: Munakata, Satoshi, et al.
Published: (2024)
by: Munakata, Satoshi, et al.
Published: (2024)
OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
DEM: Distribution Edited Model for Training with Mixed Data Distributions
by: Ram, Dhananjay, et al.
Published: (2024)
by: Ram, Dhananjay, et al.
Published: (2024)
Propaganda to Hate: A Multimodal Analysis of Arabic Memes with Multi-Agent LLMs
by: Alam, Firoj, et al.
Published: (2024)
by: Alam, Firoj, et al.
Published: (2024)
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content
by: Kmainasi, Mohamed Bayan, et al.
Published: (2024)
by: Kmainasi, Mohamed Bayan, et al.
Published: (2024)
CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
LAraBench: Benchmarking Arabic AI with Large Language Models
by: Abdelali, Ahmed, et al.
Published: (2023)
by: Abdelali, Ahmed, et al.
Published: (2023)
Similar Items
-
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025) -
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025) -
Approaching I/O-optimality for Approximate Attention
by: Papp, Pál András, et al.
Published: (2026) -
Attention Meets Reachability: Structural Equivalence and Efficiency in Grammar-Constrained LLM Decoding
by: Alpay, Faruk, et al.
Published: (2026) -
A First Runtime Analysis of the PAES-25: An Enhanced Variant of the Pareto Archived Evolution Strategy
by: Opris, Andre
Published: (2025)