Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Houjun, Murty, Shikhar, Manning, Christopher D., Csordás, Róbert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915763158253568
author Liu, Houjun
Murty, Shikhar
Manning, Christopher D.
Csordás, Róbert
author_facet Liu, Houjun
Murty, Shikhar
Manning, Christopher D.
Csordás, Róbert
contents Current approaches for scaling inference-time compute in transformers train them to emit explicit chain-of-thought tokens before producing an answer. While these methods are powerful, they are limited because they cannot be applied during pretraining and rely solely on serially-generated, natural-language verbalization. In this work, we propose Thoughtbubbles, a transformer variant that natively performs parallel adaptive computation in latent space by learning to fork or delete residual streams. Thus, tokens requiring more computation can form a "bubble" of cloned residuals in the middle of the network. Crucially, this behavior is learned during pretraining with only language modeling loss. Using half of the training budget, Thoughtbubbles outperforms the perplexity and zero-shot evals of both standard decoder LMs and those using non-adaptive parallel computation approaches. These results hold across model sizes from 150M to 1.9B. Thoughtbubbles achieves competitive GSM8K results using half of the baseline's token budget. The implicit nature of our method enables models to begin learning adaptive computation at pretraining time, paving the way to unified train-time and test-time scaling behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
Liu, Houjun
Murty, Shikhar
Manning, Christopher D.
Csordás, Róbert
Machine Learning
Artificial Intelligence
Computation and Language
Neural and Evolutionary Computing
Current approaches for scaling inference-time compute in transformers train them to emit explicit chain-of-thought tokens before producing an answer. While these methods are powerful, they are limited because they cannot be applied during pretraining and rely solely on serially-generated, natural-language verbalization. In this work, we propose Thoughtbubbles, a transformer variant that natively performs parallel adaptive computation in latent space by learning to fork or delete residual streams. Thus, tokens requiring more computation can form a "bubble" of cloned residuals in the middle of the network. Crucially, this behavior is learned during pretraining with only language modeling loss. Using half of the training budget, Thoughtbubbles outperforms the perplexity and zero-shot evals of both standard decoder LMs and those using non-adaptive parallel computation approaches. These results hold across model sizes from 150M to 1.9B. Thoughtbubbles achieves competitive GSM8K results using half of the baseline's token budget. The implicit nature of our method enables models to begin learning adaptive computation at pretraining time, paving the way to unified train-time and test-time scaling behaviors.
title Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
topic Machine Learning
Artificial Intelligence
Computation and Language
Neural and Evolutionary Computing
url https://arxiv.org/abs/2510.00219