BabyLM's First Words: Word Segmentation as a Phonological Probing Task
Fuente:
arXiv
Salvato in:
| Autori principali: | Goriely, Zébulon, Buttery, Paula |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
di: Goriely, Zébulon, et al.
Pubblicazione: (2024)
di: Goriely, Zébulon, et al.
Pubblicazione: (2024)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
di: Charpentier, Lucas, et al.
Pubblicazione: (2025)
di: Charpentier, Lucas, et al.
Pubblicazione: (2025)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
di: Choshen, Leshem, et al.
Pubblicazione: (2026)
di: Choshen, Leshem, et al.
Pubblicazione: (2026)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
What is the Best Sequence Length for BABYLM?
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
di: Trhlik, Filip, et al.
Pubblicazione: (2026)
di: Trhlik, Filip, et al.
Pubblicazione: (2026)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
di: Salhan, Suchir, et al.
Pubblicazione: (2024)
di: Salhan, Suchir, et al.
Pubblicazione: (2024)
ByteSpan: Information-Driven Subword Tokenisation
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
BabyLM's First Constructions: Causal probing provides a signal of learning
di: Rozner, Joshua, et al.
Pubblicazione: (2025)
di: Rozner, Joshua, et al.
Pubblicazione: (2025)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
di: Shen, Zhewen, et al.
Pubblicazione: (2024)
di: Shen, Zhewen, et al.
Pubblicazione: (2024)
Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
di: Martinez, Richard Diehl, et al.
Pubblicazione: (2024)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
di: Warstadt, Alex, et al.
Pubblicazione: (2025)
di: Warstadt, Alex, et al.
Pubblicazione: (2025)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
di: Haga, Akari, et al.
Pubblicazione: (2024)
di: Haga, Akari, et al.
Pubblicazione: (2024)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
di: Hu, Michael Y., et al.
Pubblicazione: (2024)
di: Hu, Michael Y., et al.
Pubblicazione: (2024)
Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
di: Padovani, Francesca, et al.
Pubblicazione: (2025)
di: Padovani, Francesca, et al.
Pubblicazione: (2025)
Are BabyLMs Second Language Learners?
di: Edman, Lukas, et al.
Pubblicazione: (2024)
di: Edman, Lukas, et al.
Pubblicazione: (2024)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
di: Zeng, Linda, et al.
Pubblicazione: (2026)
di: Zeng, Linda, et al.
Pubblicazione: (2026)
Child-directed speech facilitates production, not comprehension, in BabyLMs
di: Bunzeck, Bastian, et al.
Pubblicazione: (2026)
di: Bunzeck, Bastian, et al.
Pubblicazione: (2026)
CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
di: Capone, Luca, et al.
Pubblicazione: (2025)
di: Capone, Luca, et al.
Pubblicazione: (2025)
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
di: Edman, Lukas, et al.
Pubblicazione: (2025)
di: Edman, Lukas, et al.
Pubblicazione: (2025)
BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context
di: Matzopoulos, Alexis, et al.
Pubblicazione: (2025)
di: Matzopoulos, Alexis, et al.
Pubblicazione: (2025)
Unsupervised Classification of English Words Based on Phonological Information: Discovery of Germanic and Latinate Clusters
di: Morita, Takashi, et al.
Pubblicazione: (2025)
di: Morita, Takashi, et al.
Pubblicazione: (2025)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
di: Askari, Raha, et al.
Pubblicazione: (2025)
di: Askari, Raha, et al.
Pubblicazione: (2025)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
di: Zhang, Zihong, et al.
Pubblicazione: (2025)
di: Zhang, Zihong, et al.
Pubblicazione: (2025)
HieroLM: Egyptian Hieroglyph Recovery with Next Word Prediction Language Model
di: Cai, Xuheng, et al.
Pubblicazione: (2025)
di: Cai, Xuheng, et al.
Pubblicazione: (2025)
Using Context to Improve Word Segmentation
di: Hu, Stephanie, et al.
Pubblicazione: (2025)
di: Hu, Stephanie, et al.
Pubblicazione: (2025)
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
di: Zouhar, Vilém, et al.
Pubblicazione: (2023)
di: Zouhar, Vilém, et al.
Pubblicazione: (2023)
Parsing Through Boundaries in Chinese Word Segmentation
di: Chen, Yige, et al.
Pubblicazione: (2025)
di: Chen, Yige, et al.
Pubblicazione: (2025)
Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation
di: Wang, Xuebin, et al.
Pubblicazione: (2024)
di: Wang, Xuebin, et al.
Pubblicazione: (2024)
The LSCD Benchmark: a Testbed for Diachronic Word Meaning Tasks
di: Schlechtweg, Dominik, et al.
Pubblicazione: (2024)
di: Schlechtweg, Dominik, et al.
Pubblicazione: (2024)
Label Words as Local Task Vectors in In-Context Learning
di: Zheng, Bowen, et al.
Pubblicazione: (2024)
di: Zheng, Bowen, et al.
Pubblicazione: (2024)
Word Segmentation for Asian Languages: Chinese, Korean, and Japanese
di: Rho, Matthew, et al.
Pubblicazione: (2024)
di: Rho, Matthew, et al.
Pubblicazione: (2024)
Word Chain Generators for Prefix Normal Words
di: Adamson, Duncan, et al.
Pubblicazione: (2025)
di: Adamson, Duncan, et al.
Pubblicazione: (2025)
Probing Internal Representations of Multi-Word Verbs in Large Language Models
di: Kissane, Hassane, et al.
Pubblicazione: (2025)
di: Kissane, Hassane, et al.
Pubblicazione: (2025)
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
di: Ali, Wazir, et al.
Pubblicazione: (2024)
di: Ali, Wazir, et al.
Pubblicazione: (2024)
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
di: Jumelet, Jaap, et al.
Pubblicazione: (2025)
di: Jumelet, Jaap, et al.
Pubblicazione: (2025)
One Word Is Not Enough: Simple Prompts Improve Word Embeddings
di: Ranjan, Rajeev
Pubblicazione: (2025)
di: Ranjan, Rajeev
Pubblicazione: (2025)
Documenti analoghi
-
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
di: Goriely, Zébulon, et al.
Pubblicazione: (2025) -
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
di: Goriely, Zébulon, et al.
Pubblicazione: (2024) -
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
di: Charpentier, Lucas, et al.
Pubblicazione: (2025) -
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
di: Choshen, Leshem, et al.
Pubblicazione: (2026) -
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
di: Salhan, Suchir, et al.
Pubblicazione: (2025)