Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910260394983424 |
|---|---|
| author | Huang, Zeyi He, Xuehai Ren, LiLiang Wang, Yiping Peng, Baolin Cheng, Hao Wang, Shuohang He, Pengcheng Gao, Jianfeng Lee, Yong Jae Shen, Yelong |
| author_facet | Huang, Zeyi He, Xuehai Ren, LiLiang Wang, Yiping Peng, Baolin Cheng, Hao Wang, Shuohang He, Pengcheng Gao, Jianfeng Lee, Yong Jae Shen, Yelong |
| contents | We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this source state is already computed during ordinary decoding, LRT adds a cross-layer recurrent latent pathway across positions without inserting pause tokens or extra depth loops, and the standard attention mechanism and KV-cache interface are preserved. To pretrain this recurrence at scale without sequentially unrolling the transformer, we introduce interleaved parallel training: a single full-sequence initialization forward pass builds a shared buffer; then disjoint position subsets are refined in parallel and written back, so that all tokens receive recurrent-memory-aware supervision at roughly 2 times baseline compute. Across nanochat style backbones and a wide range of tokens-per-parameter budgets, LRT improves both language-modeling loss and in-context learning under matched effective compute while adding as little as 0.3% parameters. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_26797 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Huang, Zeyi He, Xuehai Ren, LiLiang Wang, Yiping Peng, Baolin Cheng, Hao Wang, Shuohang He, Pengcheng Gao, Jianfeng Lee, Yong Jae Shen, Yelong Machine Learning Computation and Language We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden state from the previous token as recurrent memory for the next token. Because this source state is already computed during ordinary decoding, LRT adds a cross-layer recurrent latent pathway across positions without inserting pause tokens or extra depth loops, and the standard attention mechanism and KV-cache interface are preserved. To pretrain this recurrence at scale without sequentially unrolling the transformer, we introduce interleaved parallel training: a single full-sequence initialization forward pass builds a shared buffer; then disjoint position subsets are refined in parallel and written back, so that all tokens receive recurrent-memory-aware supervision at roughly 2 times baseline compute. Across nanochat style backbones and a wide range of tokens-per-parameter budgets, LRT improves both language-modeling loss and in-context learning under matched effective compute while adding as little as 0.3% parameters. |
| title | Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2605.26797 |