Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Govindarajan, Venkata S, Rodriguez, Juan Diego, Bostrom, Kaj, Mahowald, Kyle
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915939648274432
author Govindarajan, Venkata S
Rodriguez, Juan Diego
Bostrom, Kaj
Mahowald, Kyle
author_facet Govindarajan, Venkata S
Rodriguez, Juan Diego
Bostrom, Kaj
Mahowald, Kyle
contents We present Lil-Bevo, our submission to the BabyLM Challenge. We pretrained our masked language models with three ingredients: an initial pretraining with music data, training on shorter sequences before training on longer ones, and masking specific tokens to target some of the BLiMP subtasks. Overall, our baseline models performed above chance, but far below the performance levels of larger LLMs trained on more data. We found that training on short sequences performed better than training on longer sequences.Pretraining on music may help performance marginally, but, if so, the effect seems small. Our targeted Masked Language Modeling augmentation did not seem to improve model performance in general, but did seem to help on some of the specific BLiMP tasks that we were targeting (e.g., Negative Polarity Items). Training performant LLMs on small amounts of data is a difficult but potentially informative task. While some of our techniques showed some promise, more work is needed to explore whether they can improve performance more than the modest gains here. Our code is available at https://github.com/venkatasg/Lil-Bevo and out models at https://huggingface.co/collections/venkatasg/babylm-653591cdb66f4bf68922873a
format Preprint
id arxiv_https___arxiv_org_abs_2310_17591
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
Govindarajan, Venkata S
Rodriguez, Juan Diego
Bostrom, Kaj
Mahowald, Kyle
Computation and Language
We present Lil-Bevo, our submission to the BabyLM Challenge. We pretrained our masked language models with three ingredients: an initial pretraining with music data, training on shorter sequences before training on longer ones, and masking specific tokens to target some of the BLiMP subtasks. Overall, our baseline models performed above chance, but far below the performance levels of larger LLMs trained on more data. We found that training on short sequences performed better than training on longer sequences.Pretraining on music may help performance marginally, but, if so, the effect seems small. Our targeted Masked Language Modeling augmentation did not seem to improve model performance in general, but did seem to help on some of the specific BLiMP tasks that we were targeting (e.g., Negative Polarity Items). Training performant LLMs on small amounts of data is a difficult but potentially informative task. While some of our techniques showed some promise, more work is needed to explore whether they can improve performance more than the modest gains here. Our code is available at https://github.com/venkatasg/Lil-Bevo and out models at https://huggingface.co/collections/venkatasg/babylm-653591cdb66f4bf68922873a
title Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
topic Computation and Language
url https://arxiv.org/abs/2310.17591