ByteSpan: Information-Driven Subword Tokenisation
Fuente:
arXiv
Saved in:
| Main Authors: | Goriely, Zébulon, Salhan, Suchir, Lesci, Pietro, Cheng, Julius, Buttery, Paula |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What is the Best Sequence Length for BABYLM?
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024)
by: Salhan, Suchir, et al.
Published: (2024)
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
by: Ganescu, Bianca-Mihaela, et al.
Published: (2025)
by: Ganescu, Bianca-Mihaela, et al.
Published: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
by: Martinez, Richard Diehl, et al.
Published: (2024)
by: Martinez, Richard Diehl, et al.
Published: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
by: Goriely, Zébulon, et al.
Published: (2024)
by: Goriely, Zébulon, et al.
Published: (2024)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
by: Martinez, Richard Diehl, et al.
Published: (2025)
by: Martinez, Richard Diehl, et al.
Published: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Tending Towards Stability: Convergence Challenges in Small Language Models
by: Martinez, Richard Diehl, et al.
Published: (2024)
by: Martinez, Richard Diehl, et al.
Published: (2024)
Causal Estimation of Tokenisation Bias
by: Lesci, Pietro, et al.
Published: (2025)
by: Lesci, Pietro, et al.
Published: (2025)
The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models
by: Martín, Fermín Moscoso del Prado, et al.
Published: (2026)
by: Martín, Fermín Moscoso del Prado, et al.
Published: (2026)
A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction
by: Marcheva, Mila, et al.
Published: (2026)
by: Marcheva, Mila, et al.
Published: (2026)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
by: Salhan, Suchir, et al.
Published: (2025)
by: Salhan, Suchir, et al.
Published: (2025)
Modelling the Diachronic Emergence of Phoneme Frequency Distributions
by: Martín, Fermín Moscoso del Prado, et al.
Published: (2026)
by: Martín, Fermín Moscoso del Prado, et al.
Published: (2026)
Stochasticity in Tokenisation Improves Robustness
by: Steger, Sophie, et al.
Published: (2026)
by: Steger, Sophie, et al.
Published: (2026)
Tokenisation is NP-Complete
by: Whittington, Philip, et al.
Published: (2024)
by: Whittington, Philip, et al.
Published: (2024)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
by: Gigant, Théo, et al.
Published: (2026)
by: Gigant, Théo, et al.
Published: (2026)
Tokenisation via Convex Relaxations
by: Tempus, Jan, et al.
Published: (2026)
by: Tempus, Jan, et al.
Published: (2026)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
by: Lesci, Pietro, et al.
Published: (2024)
by: Lesci, Pietro, et al.
Published: (2024)
Tokenisation over Bounded Alphabets is Hard
by: Kastreva, Violeta, et al.
Published: (2025)
by: Kastreva, Violeta, et al.
Published: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
by: Trhlik, Filip, et al.
Published: (2026)
by: Trhlik, Filip, et al.
Published: (2026)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)
by: Libovický, Jindřich, et al.
Published: (2024)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024)
by: Tsuji, Kohei, et al.
Published: (2024)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
by: Choshen, Leshem, et al.
Published: (2026)
by: Choshen, Leshem, et al.
Published: (2026)
What Language is This? Ask Your Tokenizer
by: Meister, Clara, et al.
Published: (2026)
by: Meister, Clara, et al.
Published: (2026)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
by: Meyer, Francois, et al.
Published: (2025)
by: Meyer, Francois, et al.
Published: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
Learning Mutually Informed Representations for Characters and Subwords
by: Wang, Yilin, et al.
Published: (2023)
by: Wang, Yilin, et al.
Published: (2023)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
by: Sims, Anya, et al.
Published: (2025)
by: Sims, Anya, et al.
Published: (2025)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Learning Dynamics of Meta-Learning in Small Model Pretraining
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Similar Items
-
What is the Best Sequence Length for BABYLM?
by: Salhan, Suchir, et al.
Published: (2025) -
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
by: Salhan, Suchir, et al.
Published: (2024) -
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
by: Goriely, Zébulon, et al.
Published: (2025) -
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
by: Goriely, Zébulon, et al.
Published: (2025) -
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
by: Gao, Yuan, et al.
Published: (2025)