ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: You, Haochen, Liu, Baojing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915530606116864
author You, Haochen
Liu, Baojing
author_facet You, Haochen
Liu, Baojing
contents While Transformer architectures have demonstrated impressive scalability across domains, they continue to face challenges in long-context reasoning, computational efficiency, and structural generalization - largely due to rigid layer stacking, dense attention, and reliance on positional encodings. We present ReSSFormer, a Recursive Sparse Structured Transformer that integrates three complementary innovations: Recurrent Reasoning & Memory Unit (R2MU) for iterative reasoning with bounded depth, Adaptive Sparse Attention Module (ASAM) for efficient and focused context selection, and Self-Organizing Encoder Structure (SOES) for position-free structure induction. ReSSFormer replaces conventional depth stacking with recurrent inference, substitutes full attention with token- and expert-level sparsity, and models latent token topology directly from content. Across language modeling, multi-hop QA, and structure-sensitive tasks, ReSSFormer consistently outperforms strong baselines under comparable FLOPs and parameter budgets, highlighting its scalability, efficiency, and structural flexibility.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning
You, Haochen
Liu, Baojing
Computation and Language
Networking and Internet Architecture
While Transformer architectures have demonstrated impressive scalability across domains, they continue to face challenges in long-context reasoning, computational efficiency, and structural generalization - largely due to rigid layer stacking, dense attention, and reliance on positional encodings. We present ReSSFormer, a Recursive Sparse Structured Transformer that integrates three complementary innovations: Recurrent Reasoning & Memory Unit (R2MU) for iterative reasoning with bounded depth, Adaptive Sparse Attention Module (ASAM) for efficient and focused context selection, and Self-Organizing Encoder Structure (SOES) for position-free structure induction. ReSSFormer replaces conventional depth stacking with recurrent inference, substitutes full attention with token- and expert-level sparsity, and models latent token topology directly from content. Across language modeling, multi-hop QA, and structure-sensitive tasks, ReSSFormer consistently outperforms strong baselines under comparable FLOPs and parameter budgets, highlighting its scalability, efficiency, and structural flexibility.
title ReSSFormer: A Recursive Sparse Structured Transformer for Scalable and Long-Context Reasoning
topic Computation and Language
Networking and Internet Architecture
url https://arxiv.org/abs/2510.01585