Are BabyLMs Second Language Learners?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Edman, Lukas, Bylinina, Lisa, Ghorbanpour, Faeze, Fraser, Alexander
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914993995251712
author Edman, Lukas
Bylinina, Lisa
Ghorbanpour, Faeze
Fraser, Alexander
author_facet Edman, Lukas
Bylinina, Lisa
Ghorbanpour, Faeze
Fraser, Alexander
contents This paper describes a linguistically-motivated approach to the 2024 edition of the BabyLM Challenge (Warstadt et al. 2023). Rather than pursuing a first language learning (L1) paradigm, we approach the challenge from a second language (L2) learning perspective. In L2 learning, there is a stronger focus on learning explicit linguistic information, such as grammatical notions, definitions of words or different ways of expressing a meaning. This makes L2 learning potentially more efficient and concise. We approximate this using data from Wiktionary, grammar examples either generated by an LLM or sourced from grammar books, and paraphrase data. We find that explicit information about word meaning (in our case, Wiktionary) does not boost model performance, while grammatical information can give a small improvement. The most impactful data ingredient is sentence paraphrases, with our two best models being trained on 1) a mix of paraphrase data and data from the BabyLM pretraining dataset, and 2) exclusively paraphrase data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21254
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are BabyLMs Second Language Learners?
Edman, Lukas
Bylinina, Lisa
Ghorbanpour, Faeze
Fraser, Alexander
Computation and Language
This paper describes a linguistically-motivated approach to the 2024 edition of the BabyLM Challenge (Warstadt et al. 2023). Rather than pursuing a first language learning (L1) paradigm, we approach the challenge from a second language (L2) learning perspective. In L2 learning, there is a stronger focus on learning explicit linguistic information, such as grammatical notions, definitions of words or different ways of expressing a meaning. This makes L2 learning potentially more efficient and concise. We approximate this using data from Wiktionary, grammar examples either generated by an LLM or sourced from grammar books, and paraphrase data. We find that explicit information about word meaning (in our case, Wiktionary) does not boost model performance, while grammatical information can give a small improvement. The most impactful data ingredient is sentence paraphrases, with our two best models being trained on 1) a mix of paraphrase data and data from the BabyLM pretraining dataset, and 2) exclusively paraphrase data.
title Are BabyLMs Second Language Learners?
topic Computation and Language
url https://arxiv.org/abs/2410.21254