BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912642633826304 |
|---|---|
| author | Jumelet, Jaap Fourtassi, Abdellah Haga, Akari Bunzeck, Bastian Shandilya, Bhargav Galvan-Sosa, Diana Haznitrama, Faiz Ghifari Padovani, Francesca Meyer, Francois Hu, Hai Etxaniz, Julen Prévot, Laurent He, Linyang Grandury, María Marcheva, Mila Foroutan, Negar Theodoropoulos, Nikitas Sadeghi, Pouya Song, Siyuan Salhan, Suchir Zhou, Susana Paniv, Yurii Zhang, Ziyin Bisazza, Arianna Warstadt, Alex Choshen, Leshem |
| author_facet | Jumelet, Jaap Fourtassi, Abdellah Haga, Akari Bunzeck, Bastian Shandilya, Bhargav Galvan-Sosa, Diana Haznitrama, Faiz Ghifari Padovani, Francesca Meyer, Francois Hu, Hai Etxaniz, Julen Prévot, Laurent He, Linyang Grandury, María Marcheva, Mila Foroutan, Negar Theodoropoulos, Nikitas Sadeghi, Pouya Song, Siyuan Salhan, Suchir Zhou, Susana Paniv, Yurii Zhang, Ziyin Bisazza, Arianna Warstadt, Alex Choshen, Leshem |
| contents | We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally plausible pretraining data aiming to cover the equivalent of 100M English words of content in each of 45 languages. We compile evaluation suites and train baseline models in each language. BabyBabelLM aims to facilitate multilingual pretraining and cognitive modeling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_10159 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data Jumelet, Jaap Fourtassi, Abdellah Haga, Akari Bunzeck, Bastian Shandilya, Bhargav Galvan-Sosa, Diana Haznitrama, Faiz Ghifari Padovani, Francesca Meyer, Francois Hu, Hai Etxaniz, Julen Prévot, Laurent He, Linyang Grandury, María Marcheva, Mila Foroutan, Negar Theodoropoulos, Nikitas Sadeghi, Pouya Song, Siyuan Salhan, Suchir Zhou, Susana Paniv, Yurii Zhang, Ziyin Bisazza, Arianna Warstadt, Alex Choshen, Leshem Computation and Language We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally plausible pretraining data aiming to cover the equivalent of 100M English words of content in each of 45 languages. We compile evaluation suites and train baseline models in each language. BabyBabelLM aims to facilitate multilingual pretraining and cognitive modeling. |
| title | BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.10159 |