Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zeyi, He, Xuehai, Ren, LiLiang, Wang, Yiping, Peng, Baolin, Cheng, Hao, Wang, Shuohang, He, Pengcheng, Gao, Jianfeng, Lee, Yong Jae, Shen, Yelong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910260394983424
author Huang, Zeyi
He, Xuehai
Ren, LiLiang
Wang, Yiping
Peng, Baolin
Cheng, Hao
Wang, Shuohang
He, Pengcheng
Gao, Jianfeng
Lee, Yong Jae
Shen, Yelong
author_facet Huang, Zeyi
He, Xuehai
Ren, LiLiang
Wang, Yiping
Peng, Baolin
Cheng, Hao
Wang, Shuohang
He, Pengcheng
Gao, Jianfeng
Lee, Yong Jae
Shen, Yelong
contents We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this source state is already computed during ordinary decoding, LRT adds a cross-layer recurrent latent pathway across positions without inserting pause tokens or extra depth loops, and the standard attention mechanism and KV-cache interface are preserved. To pretrain this recurrence at scale without sequentially unrolling the transformer, we introduce interleaved parallel training: a single full-sequence initialization forward pass builds a shared buffer; then disjoint position subsets are refined in parallel and written back, so that all tokens receive recurrent-memory-aware supervision at roughly 2 times baseline compute. Across nanochat style backbones and a wide range of tokens-per-parameter budgets, LRT improves both language-modeling loss and in-context learning under matched effective compute while adding as little as 0.3% parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26797
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
Huang, Zeyi
He, Xuehai
Ren, LiLiang
Wang, Yiping
Peng, Baolin
Cheng, Hao
Wang, Shuohang
He, Pengcheng
Gao, Jianfeng
Lee, Yong Jae
Shen, Yelong
Machine Learning
Computation and Language
We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this source state is already computed during ordinary decoding, LRT adds a cross-layer recurrent latent pathway across positions without inserting pause tokens or extra depth loops, and the standard attention mechanism and KV-cache interface are preserved. To pretrain this recurrence at scale without sequentially unrolling the transformer, we introduce interleaved parallel training: a single full-sequence initialization forward pass builds a shared buffer; then disjoint position subsets are refined in parallel and written back, so that all tokens receive recurrent-memory-aware supervision at roughly 2 times baseline compute. Across nanochat style backbones and a wide range of tokens-per-parameter budgets, LRT improves both language-modeling loss and in-context learning under matched effective compute while adding as little as 0.3% parameters.
title Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.26797