Saved in:
Bibliographic Details
Main Authors: Zeng, Boyi, Hao, Yiqin, Li, He, Song, Shixiang, Song, Feichen, Wang, Zitong, Huang, Siyuan, Xu, Yi, He, ZiWei, Wang, Xinbing, Lin, Zhouhan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.08220
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911500241731584
author Zeng, Boyi
Hao, Yiqin
Li, He
Song, Shixiang
Song, Feichen
Wang, Zitong
Huang, Siyuan
Xu, Yi
He, ZiWei
Wang, Xinbing
Lin, Zhouhan
author_facet Zeng, Boyi
Hao, Yiqin
Li, He
Song, Shixiang
Song, Feichen
Wang, Zitong
Huang, Siyuan
Xu, Yi
He, ZiWei
Wang, Xinbing
Lin, Zhouhan
contents Scaling large language models by increasing parameters and training data is increasingly constrained by limited high-quality corpora and rising communication costs. This work explores an alternative axis: increasing per-token computation without expanding parameters, by internalizing latent Chain-of-Thought (CoT) into pretraining. We propose Pretraining with Token-Level Adaptive Latent CoT (adaptive latent CoT), where the model generates a variable-length latent CoT trajectory before emitting each token -- allocating longer trajectories to difficult tokens and shorter (or even zero) trajectories to easy ones. Importantly, this behavior emerges naturally from one-stage pretraining on general text and reduces computation in both training and inference via token-wise adaptive halting. Experiments with Llama architectures show that adaptive latent CoT consistently improves language modeling perplexity and broad downstream accuracy, even with fewer training FLOPs than prior recurrent baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08220
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pretraining with Token-Level Adaptive Latent Chain-of-Thought
Zeng, Boyi
Hao, Yiqin
Li, He
Song, Shixiang
Song, Feichen
Wang, Zitong
Huang, Siyuan
Xu, Yi
He, ZiWei
Wang, Xinbing
Lin, Zhouhan
Computation and Language
Scaling large language models by increasing parameters and training data is increasingly constrained by limited high-quality corpora and rising communication costs. This work explores an alternative axis: increasing per-token computation without expanding parameters, by internalizing latent Chain-of-Thought (CoT) into pretraining. We propose Pretraining with Token-Level Adaptive Latent CoT (adaptive latent CoT), where the model generates a variable-length latent CoT trajectory before emitting each token -- allocating longer trajectories to difficult tokens and shorter (or even zero) trajectories to easy ones. Importantly, this behavior emerges naturally from one-stage pretraining on general text and reduces computation in both training and inference via token-wise adaptive halting. Experiments with Llama architectures show that adaptive latent CoT consistently improves language modeling perplexity and broad downstream accuracy, even with fewer training FLOPs than prior recurrent baselines.
title Pretraining with Token-Level Adaptive Latent Chain-of-Thought
topic Computation and Language
url https://arxiv.org/abs/2602.08220