MeSH: Memory-as-State-Highways for Recursive Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Chengting, Shu, Xiaobo, Wang, Yadao, Zhang, Yizhen, Wu, Haoyi, Li, Jiaang, Long, Rujiao, Chen, Ziheng, Xu, Yuchi, Su, Wenbo, Zheng, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908978995265536
author Yu, Chengting
Shu, Xiaobo
Wang, Yadao
Zhang, Yizhen
Wu, Haoyi
Li, Jiaang
Long, Rujiao
Chen, Ziheng
Xu, Yuchi
Su, Wenbo
Zheng, Bo
author_facet Yu, Chengting
Shu, Xiaobo
Wang, Yadao
Zhang, Yizhen
Wu, Haoyi
Li, Jiaang
Long, Rujiao
Chen, Ziheng
Xu, Yuchi
Su, Wenbo
Zheng, Bo
contents Recursive transformers reuse parameters and iterate over hidden states multiple times, decoupling compute depth from parameter depth. However, under matched compute, recursive models with fewer parameters often lag behind non-recursive counterparts. By probing hidden states, we trace this performance gap to two primary bottlenecks: undifferentiated computation, where the core is forced to adopt a similar computational pattern at every iteration, and information overload, where long-lived and transient information must coexist in a single hidden state. To address the issues, we introduce a Memory-as-State-Highways (MeSH) scheme, which externalizes state management into an explicit memory buffer and employs lightweight routers to dynamically diversify computation across iterations. Probing visualizations confirm that MeSH successfully resolves the pathologies by inducing functional specialization across iterations. On the Pythia suite (160M-6.9B), MeSH-enhanced recursive transformers consistently improve over recursive baselines and outperforms its larger non-recursive counterpart at the 1.4B scale, improving average downstream accuracy by +1.06% with 33% fewer non-embedding parameters. Our analysis establishes MeSH as a scalable and principled architecture for building stronger recursive models. Our code is available at https://github.com/LivingFutureLab/MeSH/ .
format Preprint
id arxiv_https___arxiv_org_abs_2510_07739
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeSH: Memory-as-State-Highways for Recursive Transformers
Yu, Chengting
Shu, Xiaobo
Wang, Yadao
Zhang, Yizhen
Wu, Haoyi
Li, Jiaang
Long, Rujiao
Chen, Ziheng
Xu, Yuchi
Su, Wenbo
Zheng, Bo
Machine Learning
Artificial Intelligence
Recursive transformers reuse parameters and iterate over hidden states multiple times, decoupling compute depth from parameter depth. However, under matched compute, recursive models with fewer parameters often lag behind non-recursive counterparts. By probing hidden states, we trace this performance gap to two primary bottlenecks: undifferentiated computation, where the core is forced to adopt a similar computational pattern at every iteration, and information overload, where long-lived and transient information must coexist in a single hidden state. To address the issues, we introduce a Memory-as-State-Highways (MeSH) scheme, which externalizes state management into an explicit memory buffer and employs lightweight routers to dynamically diversify computation across iterations. Probing visualizations confirm that MeSH successfully resolves the pathologies by inducing functional specialization across iterations. On the Pythia suite (160M-6.9B), MeSH-enhanced recursive transformers consistently improve over recursive baselines and outperforms its larger non-recursive counterpart at the 1.4B scale, improving average downstream accuracy by +1.06% with 33% fewer non-embedding parameters. Our analysis establishes MeSH as a scalable and principled architecture for building stronger recursive models. Our code is available at https://github.com/LivingFutureLab/MeSH/ .
title MeSH: Memory-as-State-Highways for Recursive Transformers
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.07739