BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Theodoropoulos, Nikitas, Filandrianos, Giorgos, Lyberatos, Vassilis, Lymperaiou, Maria, Stamou, Giorgos
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913968818225152
author Theodoropoulos, Nikitas
Filandrianos, Giorgos
Lyberatos, Vassilis
Lymperaiou, Maria
Stamou, Giorgos
author_facet Theodoropoulos, Nikitas
Filandrianos, Giorgos
Lyberatos, Vassilis
Lymperaiou, Maria
Stamou, Giorgos
contents We describe our contribution to the Strict and Strict-Small tracks of the 2nd iteration of the BabyLM Challenge. The shared task is centered around efficient pre-training given data constraints motivated by human development. In response, we study the effect of synthetic story data in language pre-training using TinyStories: a recently introduced dataset of short stories. Initially, we train GPT-Neo models on subsets of TinyStories, while varying the amount of available data. We find that, even with access to less than 100M words, the models are able to generate high-quality, original completions to a given story, and acquire substantial linguistic knowledge. To measure the effect of synthetic story data, we train LTG-BERT encoder models on a combined dataset of: a subset of TinyStories, story completions generated by GPT-Neo, and a subset of the BabyLM dataset. Our experimentation reveals that synthetic data can occasionally offer modest gains, but overall have a negative influence on linguistic understanding. Our work offers an initial study on synthesizing story data in low resource settings and underscores their potential for augmentation in data-constrained language modeling. We publicly release our models and implementation on our GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
Theodoropoulos, Nikitas
Filandrianos, Giorgos
Lyberatos, Vassilis
Lymperaiou, Maria
Stamou, Giorgos
Computation and Language
We describe our contribution to the Strict and Strict-Small tracks of the 2nd iteration of the BabyLM Challenge. The shared task is centered around efficient pre-training given data constraints motivated by human development. In response, we study the effect of synthetic story data in language pre-training using TinyStories: a recently introduced dataset of short stories. Initially, we train GPT-Neo models on subsets of TinyStories, while varying the amount of available data. We find that, even with access to less than 100M words, the models are able to generate high-quality, original completions to a given story, and acquire substantial linguistic knowledge. To measure the effect of synthetic story data, we train LTG-BERT encoder models on a combined dataset of: a subset of TinyStories, story completions generated by GPT-Neo, and a subset of the BabyLM dataset. Our experimentation reveals that synthetic data can occasionally offer modest gains, but overall have a negative influence on linguistic understanding. Our work offers an initial study on synthesizing story data in low resource settings and underscores their potential for augmentation in data-constrained language modeling. We publicly release our models and implementation on our GitHub.
title BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
topic Computation and Language
url https://arxiv.org/abs/2410.15365