Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
Fuente:
arXiv
Saved in:
| Main Authors: | Bick, Aviv, Xing, Eric, Gu, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval-Aware Distillation for Transformer-SSM Hybrids
by: Bick, Aviv, et al.
Published: (2026)
by: Bick, Aviv, et al.
Published: (2026)
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
A Comparison of Methods for Neural Network Aggregation
by: Pomerat, John, et al.
Published: (2023)
by: Pomerat, John, et al.
Published: (2023)
Recurrent Aggregators in Neural Algorithmic Reasoning
by: Xu, Kaijia, et al.
Published: (2024)
by: Xu, Kaijia, et al.
Published: (2024)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
by: Gu, Albert, et al.
Published: (2023)
by: Gu, Albert, et al.
Published: (2023)
Emergent Semantic Role Understanding in Language Models
by: Griffiths, Carla, et al.
Published: (2026)
by: Griffiths, Carla, et al.
Published: (2026)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Interpreting Affine Recurrence Learning in GPT-style Transformers
by: Bhargav, Samarth, et al.
Published: (2024)
by: Bhargav, Samarth, et al.
Published: (2024)
Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
by: Jeon, Dongjae, et al.
Published: (2024)
by: Jeon, Dongjae, et al.
Published: (2024)
Large Language Models for Missing Data Imputation: Understanding Behavior, Hallucination Effects, and Control Mechanisms
by: Mangussi, Arthur Dantas, et al.
Published: (2026)
by: Mangussi, Arthur Dantas, et al.
Published: (2026)
GATS: Gather-Attend-Scatter
by: Zolna, Konrad, et al.
Published: (2024)
by: Zolna, Konrad, et al.
Published: (2024)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
by: Kausik, Chinmaya, et al.
Published: (2026)
by: Kausik, Chinmaya, et al.
Published: (2026)
How Does Controllability Emerge In Language Models During Pretraining?
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
by: Zeng, Zhiyuan, et al.
Published: (2026)
by: Zeng, Zhiyuan, et al.
Published: (2026)
CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models
by: Lakkapragada, Venkat Akhil
Published: (2026)
by: Lakkapragada, Venkat Akhil
Published: (2026)
A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
by: Wang, Wenkai, et al.
Published: (2025)
by: Wang, Wenkai, et al.
Published: (2025)
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
by: Gu, Shangding
Published: (2026)
by: Gu, Shangding
Published: (2026)
Towards Understanding Extrapolation: a Causal Lens
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
by: Aviv, Gilad, et al.
Published: (2025)
by: Aviv, Gilad, et al.
Published: (2025)
Meta Reinforcement Learning with Finite Training Tasks -- a Density Estimation Approach
by: Rimon, Zohar, et al.
Published: (2022)
by: Rimon, Zohar, et al.
Published: (2022)
Improved Generalization of Weight Space Networks via Augmentations
by: Shamsian, Aviv, et al.
Published: (2024)
by: Shamsian, Aviv, et al.
Published: (2024)
Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation Tree
by: Feng, Lang, et al.
Published: (2024)
by: Feng, Lang, et al.
Published: (2024)
Chimera: State Space Models Beyond Sequences
by: Lahoti, Aakash, et al.
Published: (2025)
by: Lahoti, Aakash, et al.
Published: (2025)
Reviewing AI's Role in Non-Muscle-Invasive Bladder Cancer Recurrence Prediction
by: Abbas, Saram, et al.
Published: (2024)
by: Abbas, Saram, et al.
Published: (2024)
Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models
by: Yang, Xiao-Wen, et al.
Published: (2026)
by: Yang, Xiao-Wen, et al.
Published: (2026)
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
by: Hwang, Sukjun, et al.
Published: (2024)
by: Hwang, Sukjun, et al.
Published: (2024)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
by: Sieber, Jerome, et al.
Published: (2024)
by: Sieber, Jerome, et al.
Published: (2024)
Understanding Dynamic Compute Allocation in Recurrent Transformers
by: Moosa, Ibraheem Muhammad, et al.
Published: (2026)
by: Moosa, Ibraheem Muhammad, et al.
Published: (2026)
On the Benefits of Memory for Modeling Time-Dependent PDEs
by: Ruiz, Ricardo Buitrago, et al.
Published: (2024)
by: Ruiz, Ricardo Buitrago, et al.
Published: (2024)
Understanding the Functional Roles of Modelling Components in Spiking Neural Networks
by: Yin, Huifeng, et al.
Published: (2024)
by: Yin, Huifeng, et al.
Published: (2024)
Symbol-Equivariant Recurrent Reasoning Models
by: Freinschlag, Richard, et al.
Published: (2026)
by: Freinschlag, Richard, et al.
Published: (2026)
Understanding the Mechanisms of Fast Hyperparameter Transfer
by: Ghosh, Nikhil, et al.
Published: (2025)
by: Ghosh, Nikhil, et al.
Published: (2025)
Convergence Analysis of Aggregation-Broadcast in LoRA-enabled Distributed Fine-Tuning
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Deep Reinforcement Multi-agent Learning framework for Information Gathering with Local Gaussian Processes for Water Monitoring
by: Luis, Samuel Yanes, et al.
Published: (2024)
by: Luis, Samuel Yanes, et al.
Published: (2024)
Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
by: Sharifi, Mohammadreza, et al.
Published: (2025)
by: Sharifi, Mohammadreza, et al.
Published: (2025)
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
by: Garg, Anisha, et al.
Published: (2025)
by: Garg, Anisha, et al.
Published: (2025)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
by: Shen, Jucheng, et al.
Published: (2026)
by: Shen, Jucheng, et al.
Published: (2026)
Similar Items
-
Retrieval-Aware Distillation for Transformer-SSM Hybrids
by: Bick, Aviv, et al.
Published: (2026) -
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025) -
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024) -
A Comparison of Methods for Neural Network Aggregation
by: Pomerat, John, et al.
Published: (2023) -
Recurrent Aggregators in Neural Algorithmic Reasoning
by: Xu, Kaijia, et al.
Published: (2024)