ByteSpan: Information-Driven Subword Tokenisation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goriely, Zébulon, Salhan, Suchir, Lesci, Pietro, Cheng, Julius, Buttery, Paula |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What is the Best Sequence Length for BABYLM?
von: Salhan, Suchir, et al.
Veröffentlicht: (2025)
von: Salhan, Suchir, et al.
Veröffentlicht: (2025)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
von: Salhan, Suchir, et al.
Veröffentlicht: (2024)
von: Salhan, Suchir, et al.
Veröffentlicht: (2024)
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling
von: Ganescu, Bianca-Mihaela, et al.
Veröffentlicht: (2025)
von: Ganescu, Bianca-Mihaela, et al.
Veröffentlicht: (2025)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2024)
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
von: Goriely, Zébulon, et al.
Veröffentlicht: (2024)
von: Goriely, Zébulon, et al.
Veröffentlicht: (2024)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2025)
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
Tending Towards Stability: Convergence Challenges in Small Language Models
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2024)
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2024)
Causal Estimation of Tokenisation Bias
von: Lesci, Pietro, et al.
Veröffentlicht: (2025)
von: Lesci, Pietro, et al.
Veröffentlicht: (2025)
The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models
von: Martín, Fermín Moscoso del Prado, et al.
Veröffentlicht: (2026)
von: Martín, Fermín Moscoso del Prado, et al.
Veröffentlicht: (2026)
A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction
von: Marcheva, Mila, et al.
Veröffentlicht: (2026)
von: Marcheva, Mila, et al.
Veröffentlicht: (2026)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
von: Salhan, Suchir, et al.
Veröffentlicht: (2025)
von: Salhan, Suchir, et al.
Veröffentlicht: (2025)
Modelling the Diachronic Emergence of Phoneme Frequency Distributions
von: Martín, Fermín Moscoso del Prado, et al.
Veröffentlicht: (2026)
von: Martín, Fermín Moscoso del Prado, et al.
Veröffentlicht: (2026)
Stochasticity in Tokenisation Improves Robustness
von: Steger, Sophie, et al.
Veröffentlicht: (2026)
von: Steger, Sophie, et al.
Veröffentlicht: (2026)
Tokenisation is NP-Complete
von: Whittington, Philip, et al.
Veröffentlicht: (2024)
von: Whittington, Philip, et al.
Veröffentlicht: (2024)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
von: Gigant, Théo, et al.
Veröffentlicht: (2026)
von: Gigant, Théo, et al.
Veröffentlicht: (2026)
Tokenisation via Convex Relaxations
von: Tempus, Jan, et al.
Veröffentlicht: (2026)
von: Tempus, Jan, et al.
Veröffentlicht: (2026)
AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
von: Lesci, Pietro, et al.
Veröffentlicht: (2024)
Tokenisation over Bounded Alphabets is Hard
von: Kastreva, Violeta, et al.
Veröffentlicht: (2025)
von: Kastreva, Violeta, et al.
Veröffentlicht: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
von: Trhlik, Filip, et al.
Veröffentlicht: (2026)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
von: Batsuren, Khuyagbaatar, et al.
Veröffentlicht: (2024)
Distributional Properties of Subword Regularization
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Lexically Grounded Subword Segmentation
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024)
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
von: Tsuji, Kohei, et al.
Veröffentlicht: (2024)
von: Tsuji, Kohei, et al.
Veröffentlicht: (2024)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
von: Choshen, Leshem, et al.
Veröffentlicht: (2026)
von: Choshen, Leshem, et al.
Veröffentlicht: (2026)
What Language is This? Ask Your Tokenizer
von: Meister, Clara, et al.
Veröffentlicht: (2026)
von: Meister, Clara, et al.
Veröffentlicht: (2026)
Subword Tokenization Strategies for Kurdish Word Embeddings
von: Salehi, Ali, et al.
Veröffentlicht: (2025)
von: Salehi, Ali, et al.
Veröffentlicht: (2025)
Subword models struggle with word learning, but surprisal hides it
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
von: Meyer, Francois, et al.
Veröffentlicht: (2025)
von: Meyer, Francois, et al.
Veröffentlicht: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
von: Parra, Iñigo
Veröffentlicht: (2024)
von: Parra, Iñigo
Veröffentlicht: (2024)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
Learning Mutually Informed Representations for Characters and Subwords
von: Wang, Yilin, et al.
Veröffentlicht: (2023)
von: Wang, Yilin, et al.
Veröffentlicht: (2023)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
von: Sims, Anya, et al.
Veröffentlicht: (2025)
von: Sims, Anya, et al.
Veröffentlicht: (2025)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
von: Zouhar, Vilém
Veröffentlicht: (2024)
von: Zouhar, Vilém
Veröffentlicht: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
Learning Dynamics of Meta-Learning in Small Model Pretraining
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
von: Africa, David Demitri, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What is the Best Sequence Length for BABYLM?
von: Salhan, Suchir, et al.
Veröffentlicht: (2025) -
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
von: Salhan, Suchir, et al.
Veröffentlicht: (2024) -
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025) -
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025) -
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
von: Gao, Yuan, et al.
Veröffentlicht: (2025)