Mask and You Shall Receive: Optimizing Masked Language Modeling For Pretraining BabyLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Edman, Lukas, Fraser, Alexander |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are BabyLMs Second Language Learners?
di: Edman, Lukas, et al.
Pubblicazione: (2024)
di: Edman, Lukas, et al.
Pubblicazione: (2024)
Child-directed speech facilitates production, not comprehension, in BabyLMs
di: Bunzeck, Bastian, et al.
Pubblicazione: (2026)
di: Bunzeck, Bastian, et al.
Pubblicazione: (2026)
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context
di: Matzopoulos, Alexis, et al.
Pubblicazione: (2025)
di: Matzopoulos, Alexis, et al.
Pubblicazione: (2025)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
di: Askari, Raha, et al.
Pubblicazione: (2025)
di: Askari, Raha, et al.
Pubblicazione: (2025)
CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
di: Capone, Luca, et al.
Pubblicazione: (2025)
di: Capone, Luca, et al.
Pubblicazione: (2025)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
di: Charpentier, Lucas, et al.
Pubblicazione: (2025)
di: Charpentier, Lucas, et al.
Pubblicazione: (2025)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
di: Shen, Zhewen, et al.
Pubblicazione: (2024)
di: Shen, Zhewen, et al.
Pubblicazione: (2024)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
di: Choshen, Leshem, et al.
Pubblicazione: (2026)
di: Choshen, Leshem, et al.
Pubblicazione: (2026)
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing
di: Trhlik, Filip, et al.
Pubblicazione: (2026)
di: Trhlik, Filip, et al.
Pubblicazione: (2026)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
di: Warstadt, Alex, et al.
Pubblicazione: (2025)
di: Warstadt, Alex, et al.
Pubblicazione: (2025)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
di: Hu, Michael Y., et al.
Pubblicazione: (2024)
di: Hu, Michael Y., et al.
Pubblicazione: (2024)
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
di: Edman, Lukas, et al.
Pubblicazione: (2025)
di: Edman, Lukas, et al.
Pubblicazione: (2025)
CUTE: Measuring LLMs' Understanding of Their Tokens
di: Edman, Lukas, et al.
Pubblicazione: (2024)
di: Edman, Lukas, et al.
Pubblicazione: (2024)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
di: Haga, Akari, et al.
Pubblicazione: (2024)
di: Haga, Akari, et al.
Pubblicazione: (2024)
BabyLM's First Constructions: Causal probing provides a signal of learning
di: Rozner, Joshua, et al.
Pubblicazione: (2025)
di: Rozner, Joshua, et al.
Pubblicazione: (2025)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
Should We Still Pretrain Encoders with Masked Language Modeling?
di: Gisserot-Boukhlef, Hippolyte, et al.
Pubblicazione: (2025)
di: Gisserot-Boukhlef, Hippolyte, et al.
Pubblicazione: (2025)
Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
di: Padovani, Francesca, et al.
Pubblicazione: (2025)
di: Padovani, Francesca, et al.
Pubblicazione: (2025)
Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
di: Salhan, Suchir, et al.
Pubblicazione: (2025)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
di: Zeng, Linda, et al.
Pubblicazione: (2026)
di: Zeng, Linda, et al.
Pubblicazione: (2026)
A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs
di: Feldman, Anna, et al.
Pubblicazione: (2026)
di: Feldman, Anna, et al.
Pubblicazione: (2026)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs
di: Yu, Longxuan, et al.
Pubblicazione: (2026)
di: Yu, Longxuan, et al.
Pubblicazione: (2026)
Entropy-aware Masking for Masked Language Modeling
di: Srinivasagan, Gokul, et al.
Pubblicazione: (2026)
di: Srinivasagan, Gokul, et al.
Pubblicazione: (2026)
You Shall Know a Tool by the Traces it Leaves: The Predictability of Sentiment Analysis Tools
di: Baumartz, Daniel, et al.
Pubblicazione: (2024)
di: Baumartz, Daniel, et al.
Pubblicazione: (2024)
Dynamic Masking Rate Schedules for MLM Pretraining
di: Ankner, Zachary, et al.
Pubblicazione: (2023)
di: Ankner, Zachary, et al.
Pubblicazione: (2023)
MaskTab: Scalable Masked Tabular Pretraining with Scaling Laws and Distillation for Industrial Classification
di: Zheng, Bo, et al.
Pubblicazione: (2026)
di: Zheng, Bo, et al.
Pubblicazione: (2026)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
di: Bhaskar, Adithya, et al.
Pubblicazione: (2025)
di: Bhaskar, Adithya, et al.
Pubblicazione: (2025)
Inconsistencies in Masked Language Models
di: Young, Tom, et al.
Pubblicazione: (2022)
di: Young, Tom, et al.
Pubblicazione: (2022)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
di: Oba, Daisuke, et al.
Pubblicazione: (2026)
di: Oba, Daisuke, et al.
Pubblicazione: (2026)
MaskLID: Code-Switching Language Identification through Iterative Masking
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2024)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2024)
Activation Steering for Masked Diffusion Language Models
di: Shnaidman, Adi, et al.
Pubblicazione: (2025)
di: Shnaidman, Adi, et al.
Pubblicazione: (2025)
Are You There God? Lightweight Narrative Annotation of Christian Fiction with LMs
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2025)
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2025)
Faithfulness Measurable Masked Language Models
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
Representation Deficiency in Masked Language Modeling
di: Meng, Yu, et al.
Pubblicazione: (2023)
di: Meng, Yu, et al.
Pubblicazione: (2023)
Simple and Effective Masked Diffusion Language Models
di: Sahoo, Subham Sekhar, et al.
Pubblicazione: (2024)
di: Sahoo, Subham Sekhar, et al.
Pubblicazione: (2024)
Masked Diffusion Language Models with Frequency-Informed Training
di: Kosmopoulou, Despoina, et al.
Pubblicazione: (2025)
di: Kosmopoulou, Despoina, et al.
Pubblicazione: (2025)
AntLM: Bridging Causal and Masked Language Models
di: Yu, Xinru, et al.
Pubblicazione: (2024)
di: Yu, Xinru, et al.
Pubblicazione: (2024)
Parallel Corpus Augmentation using Masked Language Models
di: Kumari, Vibhuti, et al.
Pubblicazione: (2024)
di: Kumari, Vibhuti, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Are BabyLMs Second Language Learners?
di: Edman, Lukas, et al.
Pubblicazione: (2024) -
Child-directed speech facilitates production, not comprehension, in BabyLMs
di: Bunzeck, Bastian, et al.
Pubblicazione: (2026) -
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025) -
BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context
di: Matzopoulos, Alexis, et al.
Pubblicazione: (2025) -
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
di: Askari, Raha, et al.
Pubblicazione: (2025)