Synthetic bootstrapped pretraining

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Zitong, Zhang, Aonan, Liu, Hong, Hashimoto, Tatsunori, Candès, Emmanuel, Wang, Chong, Pang, Ruoming
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914199248044032
author Yang, Zitong
Zhang, Aonan
Liu, Hong
Hashimoto, Tatsunori
Candès, Emmanuel
Wang, Chong
Pang, Ruoming
author_facet Yang, Zitong
Zhang, Aonan
Liu, Hong
Hashimoto, Tatsunori
Candès, Emmanuel
Wang, Chong
Pang, Ruoming
contents We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dataset and then leverages it to synthesize a vast new corpus for joint training. While the standard pretraining teaches LMs to learn causal correlations among tokens within a single document, it is not designed to efficiently model the rich, learnable inter-document correlations that can potentially lead to better performance. We validate SBP by designing a compute-matched pretraining setup and pretrain a 3B-parameter and a 6B-parameter model on up to 1T tokens from scratch. We find SBP consistently improves upon a strong repetition baseline and delivers up to 60% of performance improvement attainable by an oracle upper bound with access to 20x more unique data. Qualitative analysis reveals that the synthesized documents go beyond mere paraphrases -- SBP first abstracts a core concept from the seed material and then crafts a new narration on top of it. Besides strong empirical performance, SBP admits a natural Bayesian interpretation: the synthesizer implicitly learns to abstract the latent concepts shared between related documents.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthetic bootstrapped pretraining
Yang, Zitong
Zhang, Aonan
Liu, Hong
Hashimoto, Tatsunori
Candès, Emmanuel
Wang, Chong
Pang, Ruoming
Computation and Language
Artificial Intelligence
We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dataset and then leverages it to synthesize a vast new corpus for joint training. While the standard pretraining teaches LMs to learn causal correlations among tokens within a single document, it is not designed to efficiently model the rich, learnable inter-document correlations that can potentially lead to better performance. We validate SBP by designing a compute-matched pretraining setup and pretrain a 3B-parameter and a 6B-parameter model on up to 1T tokens from scratch. We find SBP consistently improves upon a strong repetition baseline and delivers up to 60% of performance improvement attainable by an oracle upper bound with access to 20x more unique data. Qualitative analysis reveals that the synthesized documents go beyond mere paraphrases -- SBP first abstracts a core concept from the seed material and then crafts a new narration on top of it. Besides strong empirical performance, SBP admits a natural Bayesian interpretation: the synthesizer implicitly learns to abstract the latent concepts shared between related documents.
title Synthetic bootstrapped pretraining
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.15248