Mini Minds: Exploring Bebeshka and Zlata Baby Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Proskurina, Irina, Metzler, Guillaume, Velcin, Julien
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929332529659904
author Proskurina, Irina
Metzler, Guillaume
Velcin, Julien
author_facet Proskurina, Irina
Metzler, Guillaume
Velcin, Julien
contents In this paper, we describe the University of Lyon 2 submission to the Strict-Small track of the BabyLM competition. The shared task is created with an emphasis on small-scale language modelling from scratch on limited-size data and human language acquisition. Dataset released for the Strict-Small track has 10M words, which is comparable to children's vocabulary size. We approach the task with an architecture search, minimizing masked language modelling loss on the data of the shared task. Having found an optimal configuration, we introduce two small-size language models (LMs) that were submitted for evaluation, a 4-layer encoder with 8 attention heads and a 6-layer decoder model with 12 heads which we term Bebeshka and Zlata, respectively. Despite being half the scale of the baseline LMs, our proposed models achieve comparable performance. We further explore the applicability of small-scale language models in tasks involving moral judgment, aligning their predictions with human values. These findings highlight the potential of compact LMs in addressing practical language understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2311_03216
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mini Minds: Exploring Bebeshka and Zlata Baby Models
Proskurina, Irina
Metzler, Guillaume
Velcin, Julien
Computation and Language
Artificial Intelligence
Machine Learning
In this paper, we describe the University of Lyon 2 submission to the Strict-Small track of the BabyLM competition. The shared task is created with an emphasis on small-scale language modelling from scratch on limited-size data and human language acquisition. Dataset released for the Strict-Small track has 10M words, which is comparable to children's vocabulary size. We approach the task with an architecture search, minimizing masked language modelling loss on the data of the shared task. Having found an optimal configuration, we introduce two small-size language models (LMs) that were submitted for evaluation, a 4-layer encoder with 8 attention heads and a 6-layer decoder model with 12 heads which we term Bebeshka and Zlata, respectively. Despite being half the scale of the baseline LMs, our proposed models achieve comparable performance. We further explore the applicability of small-scale language models in tasks involving moral judgment, aligning their predictions with human values. These findings highlight the potential of compact LMs in addressing practical language understanding tasks.
title Mini Minds: Exploring Bebeshka and Zlata Baby Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2311.03216