Which Pieces Does Unigram Tokenization Really Need?
Fuente:
arXiv
Saved in:
| Main Authors: | Land, Sander, Pinter, Yuval |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)
by: Land, Sander, et al.
Published: (2024)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025)
by: Schmidt, Craig W., et al.
Published: (2025)
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023)
by: Gao, Yifei, et al.
Published: (2023)
Does It Make Sense to Explain a Black Box With Another Black Box?
by: Delaunay, Julien, et al.
Published: (2024)
by: Delaunay, Julien, et al.
Published: (2024)
Fast Quiet-STaR: Thinking Without Thought Tokens
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
by: Yao, Huaiyuan, et al.
Published: (2026)
by: Yao, Huaiyuan, et al.
Published: (2026)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
by: Ikoma, Hayato, et al.
Published: (2025)
by: Ikoma, Hayato, et al.
Published: (2025)
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
by: Alkan, Atilla Kaan, et al.
Published: (2025)
by: Alkan, Atilla Kaan, et al.
Published: (2025)
Duluth at SemEval-2025 Task 7: TF-IDF with Optimized Vector Dimensions for Multilingual Fact-Checked Claim Retrieval
by: Syed, Shujauddin, et al.
Published: (2025)
by: Syed, Shujauddin, et al.
Published: (2025)
IteRABRe: Iterative Recovery-Aided Block Reduction
by: Wibowo, Haryo Akbarianto, et al.
Published: (2025)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations
by: Bhattacharya, Shamik, et al.
Published: (2026)
by: Bhattacharya, Shamik, et al.
Published: (2026)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2024)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
by: Li, Yongchang, et al.
Published: (2024)
by: Li, Yongchang, et al.
Published: (2024)
BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
by: Wibowo, Haryo Akbarianto, et al.
Published: (2026)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2026)
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
by: Johnson, Warren
Published: (2026)
by: Johnson, Warren
Published: (2026)
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
by: Liu, Aiwei, et al.
Published: (2024)
by: Liu, Aiwei, et al.
Published: (2024)
Surfing the modeling of PoS taggers in low-resource scenarios
by: Ferro, Manuel Vilares, et al.
Published: (2024)
by: Ferro, Manuel Vilares, et al.
Published: (2024)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
TREX: Tokenizer Regression for Optimal Data Mixture
by: Won, Inho, et al.
Published: (2026)
by: Won, Inho, et al.
Published: (2026)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
by: Souani, Badr, et al.
Published: (2025)
by: Souani, Badr, et al.
Published: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
by: Li, Mingda, et al.
Published: (2024)
by: Li, Mingda, et al.
Published: (2024)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
A Legal Framework for Natural Language Processing Model Training in Portugal
by: Almeida, Rúben, et al.
Published: (2024)
by: Almeida, Rúben, et al.
Published: (2024)
EvidenceMap: Learning Evidence Analysis to Unleash the Power of Small Language Models for Biomedical Question Answering
by: Zong, Chang, et al.
Published: (2025)
by: Zong, Chang, et al.
Published: (2025)
Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs
by: Rakshit, Supantho, et al.
Published: (2025)
by: Rakshit, Supantho, et al.
Published: (2025)
T-VEC: A Telecom-Specific Vectorization Model with Enhanced Semantic Understanding via Deep Triplet Loss Fine-Tuning
by: Ethiraj, Vignesh, et al.
Published: (2025)
by: Ethiraj, Vignesh, et al.
Published: (2025)
From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation
by: Oduwole, Mardiyyah, et al.
Published: (2025)
by: Oduwole, Mardiyyah, et al.
Published: (2025)
Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models
by: Laurence, Timothy, et al.
Published: (2025)
by: Laurence, Timothy, et al.
Published: (2025)
Auto prompt sql: a resource-efficient architecture for text-to-sql translation in constrained environments
by: Tang, Zetong, et al.
Published: (2025)
by: Tang, Zetong, et al.
Published: (2025)
CausalSent: Interpretable Sentiment Classification with RieszNet
by: Frees, Daniel, et al.
Published: (2025)
by: Frees, Daniel, et al.
Published: (2025)
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
by: Harris, Joshua, et al.
Published: (2025)
by: Harris, Joshua, et al.
Published: (2025)
Similar Items
-
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024) -
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024) -
Boundless Byte Pair Encoding: Breaking the Pre-tokenization Barrier
by: Schmidt, Craig W., et al.
Published: (2025) -
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023) -
Does It Make Sense to Explain a Black Box With Another Black Box?
by: Delaunay, Julien, et al.
Published: (2024)