Understanding and Improving Length Generalization in Recurrent Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ruiz, Ricardo Buitrago, Gu, Albert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Benefits of Memory for Modeling Time-Dependent PDEs
von: Ruiz, Ricardo Buitrago, et al.
Veröffentlicht: (2024)
von: Ruiz, Ricardo Buitrago, et al.
Veröffentlicht: (2024)
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
von: Bick, Aviv, et al.
Veröffentlicht: (2025)
von: Bick, Aviv, et al.
Veröffentlicht: (2025)
Length Generalization with Log-Depth Recurrent Units
von: Pert, Charles, et al.
Veröffentlicht: (2026)
von: Pert, Charles, et al.
Veröffentlicht: (2026)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025)
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
von: Cheng, Zicong, et al.
Veröffentlicht: (2026)
von: Cheng, Zicong, et al.
Veröffentlicht: (2026)
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
von: Bick, Aviv, et al.
Veröffentlicht: (2025)
von: Bick, Aviv, et al.
Veröffentlicht: (2025)
A Formal Framework for Understanding Length Generalization in Transformers
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
von: Dao, Tri, et al.
Veröffentlicht: (2024)
von: Dao, Tri, et al.
Veröffentlicht: (2024)
Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
von: Jeon, Dongjae, et al.
Veröffentlicht: (2024)
von: Jeon, Dongjae, et al.
Veröffentlicht: (2024)
Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks
von: Breslow, Nathan
Veröffentlicht: (2025)
von: Breslow, Nathan
Veröffentlicht: (2025)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
Predicting the Stay Length of Patients in Hospitals using Convolutional Gated Recurrent Deep Learning Model
von: Neshat, Mehdi, et al.
Veröffentlicht: (2024)
von: Neshat, Mehdi, et al.
Veröffentlicht: (2024)
Non-Asymptotic Length Generalization
von: Chen, Thomas, et al.
Veröffentlicht: (2025)
von: Chen, Thomas, et al.
Veröffentlicht: (2025)
Looped Transformers for Length Generalization
von: Fan, Ying, et al.
Veröffentlicht: (2024)
von: Fan, Ying, et al.
Veröffentlicht: (2024)
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
von: Hwang, Sukjun, et al.
Veröffentlicht: (2024)
von: Hwang, Sukjun, et al.
Veröffentlicht: (2024)
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
von: Hwang, Sukjun, et al.
Veröffentlicht: (2025)
von: Hwang, Sukjun, et al.
Veröffentlicht: (2025)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
von: Gu, Albert, et al.
Veröffentlicht: (2023)
von: Gu, Albert, et al.
Veröffentlicht: (2023)
Quantitative Bounds for Length Generalization in Transformers
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
Mamba-3: Improved Sequence Modeling using State Space Principles
von: Lahoti, Aakash, et al.
Veröffentlicht: (2026)
von: Lahoti, Aakash, et al.
Veröffentlicht: (2026)
Adaptive Length Image Tokenization via Recurrent Allocation
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
von: Duggal, Shivam, et al.
Veröffentlicht: (2024)
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
von: Malach, Eran, et al.
Veröffentlicht: (2025)
von: Malach, Eran, et al.
Veröffentlicht: (2025)
ARC-AGI Without Pretraining
von: Liao, Isaac, et al.
Veröffentlicht: (2025)
von: Liao, Isaac, et al.
Veröffentlicht: (2025)
Interpreting Affine Recurrence Learning in GPT-style Transformers
von: Bhargav, Samarth, et al.
Veröffentlicht: (2024)
von: Bhargav, Samarth, et al.
Veröffentlicht: (2024)
Learning Variable-Length Tokenization for Generative Recommendation
von: Wang, Minhao, et al.
Veröffentlicht: (2026)
von: Wang, Minhao, et al.
Veröffentlicht: (2026)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
On Provable Length and Compositional Generalization
von: Ahuja, Kartik, et al.
Veröffentlicht: (2024)
von: Ahuja, Kartik, et al.
Veröffentlicht: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Fragility-aware Classification for Understanding Risk and Improving Generalization
von: Yang, Chen, et al.
Veröffentlicht: (2025)
von: Yang, Chen, et al.
Veröffentlicht: (2025)
Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Generalization and Risk Bounds for Recurrent Neural Networks
von: Cheng, Xuewei, et al.
Veröffentlicht: (2024)
von: Cheng, Xuewei, et al.
Veröffentlicht: (2024)
Probing Length Generalization in Mamba via Image Reconstruction
von: Rathjens, Jan, et al.
Veröffentlicht: (2026)
von: Rathjens, Jan, et al.
Veröffentlicht: (2026)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
von: Sieber, Jerome, et al.
Veröffentlicht: (2024)
von: Sieber, Jerome, et al.
Veröffentlicht: (2024)
A Theory of Machine Understanding via the Minimum Description Length Principle
von: Zhang, Canlin, et al.
Veröffentlicht: (2025)
von: Zhang, Canlin, et al.
Veröffentlicht: (2025)
Uncertainty-Gated Generative Modeling
von: Gu, Xingrui, et al.
Veröffentlicht: (2026)
von: Gu, Xingrui, et al.
Veröffentlicht: (2026)
Understanding Dynamic Compute Allocation in Recurrent Transformers
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes
von: Arora, Ashok, et al.
Veröffentlicht: (2025)
von: Arora, Ashok, et al.
Veröffentlicht: (2025)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
von: De, Soham, et al.
Veröffentlicht: (2024)
von: De, Soham, et al.
Veröffentlicht: (2024)
On Vanishing Variance in Transformer Length Generalization
von: Li, Ruining, et al.
Veröffentlicht: (2025)
von: Li, Ruining, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Benefits of Memory for Modeling Time-Dependent PDEs
von: Ruiz, Ricardo Buitrago, et al.
Veröffentlicht: (2024) -
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
von: Bick, Aviv, et al.
Veröffentlicht: (2025) -
Length Generalization with Log-Depth Recurrent Units
von: Pert, Charles, et al.
Veröffentlicht: (2026) -
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
von: Leng, Jiaqi, et al.
Veröffentlicht: (2025) -
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
von: Cheng, Zicong, et al.
Veröffentlicht: (2026)