PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zeng, Boyi, Li, He, Song, Shixiang, Wang, Yixuan, Wang, Zitong, He, Ziwei, Wang, Xinbing, Lin, Zhouhan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910044483747840
author Zeng, Boyi
Li, He
Song, Shixiang
Wang, Yixuan
Wang, Zitong
He, Ziwei
Wang, Xinbing
Lin, Zhouhan
author_facet Zeng, Boyi
Li, He
Song, Shixiang
Wang, Yixuan
Wang, Zitong
He, Ziwei
Wang, Xinbing
Lin, Zhouhan
contents The remarkable success of Chain-of-Thought (CoT), which enhances performance by scaling generation steps at test-time, inspires us to ask: can we leverage a similar scaling of computational steps during pretraining to improve the generation of each individual token? To address this, we propose a novel pre-training methodology: Pretraining Language Models with Latent Thoughts (PonderLM-2). Our approach pretrains a language model (LM) to first generate an intermediate latent thought-the last hidden state of the current position-which is then used as input to predict the actual subsequent token. This additional computational step enables the LM to refine its prediction within unconstrained continuous space. Our experiments demonstrate that, at an identical inference cost, a LM that generates one additional latent thought per token outperforms a standard model with double the parameters. For instance, our PonderLM-2-Pythia-1.4B, pretrained on 300B tokens from the Pile, significantly surpasses the vanilla Pythia-2.8B trained on the same data on both language modeling and a range of general downstream tasks. Furthermore, increasing the number of latent thoughts generated before each actual token-forming a chain analogous to CoT-consistently improves the model's performance. The code is available at https://github.com/LUMIA-Group/PonderLM-2.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
Zeng, Boyi
Li, He
Song, Shixiang
Wang, Yixuan
Wang, Zitong
He, Ziwei
Wang, Xinbing
Lin, Zhouhan
Computation and Language
The remarkable success of Chain-of-Thought (CoT), which enhances performance by scaling generation steps at test-time, inspires us to ask: can we leverage a similar scaling of computational steps during pretraining to improve the generation of each individual token? To address this, we propose a novel pre-training methodology: Pretraining Language Models with Latent Thoughts (PonderLM-2). Our approach pretrains a language model (LM) to first generate an intermediate latent thought-the last hidden state of the current position-which is then used as input to predict the actual subsequent token. This additional computational step enables the LM to refine its prediction within unconstrained continuous space. Our experiments demonstrate that, at an identical inference cost, a LM that generates one additional latent thought per token outperforms a standard model with double the parameters. For instance, our PonderLM-2-Pythia-1.4B, pretrained on 300B tokens from the Pile, significantly surpasses the vanilla Pythia-2.8B trained on the same data on both language modeling and a range of general downstream tasks. Furthermore, increasing the number of latent thoughts generated before each actual token-forming a chain analogous to CoT-consistently improves the model's performance. The code is available at https://github.com/LUMIA-Group/PonderLM-2.
title PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
topic Computation and Language
url https://arxiv.org/abs/2509.23184