StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shida, Li, Qianxiao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongSSM: On the Length Extension of State-space Models in Language Modelling
by: Wang, Shida
Published: (2024)
by: Wang, Shida
Published: (2024)
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Hybrid Energy-Based Models for Physical AI: Provably Stable Identification of Port-Hamiltonian Dynamics
by: Betteti, Simone, et al.
Published: (2026)
by: Betteti, Simone, et al.
Published: (2026)
Learning Macroscopic Dynamics from Partial Microscopic Observations
by: Chen, Mengyi, et al.
Published: (2024)
by: Chen, Mengyi, et al.
Published: (2024)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)
by: Hoover, Benjamin, et al.
Published: (2023)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
by: Kitouni, Ouail, et al.
Published: (2024)
by: Kitouni, Ouail, et al.
Published: (2024)
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
An Analysis and Mitigation of the Reversal Curse
by: Lv, Ang, et al.
Published: (2023)
by: Lv, Ang, et al.
Published: (2023)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
by: Rezazadeh, Alireza, et al.
Published: (2025)
by: Rezazadeh, Alireza, et al.
Published: (2025)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
by: Guo, Qingyan, et al.
Published: (2024)
by: Guo, Qingyan, et al.
Published: (2024)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
by: Kapl, Ferdinand, et al.
Published: (2025)
by: Kapl, Ferdinand, et al.
Published: (2025)
Memp: Exploring Agent Procedural Memory
by: Fang, Runnan, et al.
Published: (2025)
by: Fang, Runnan, et al.
Published: (2025)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
by: Luo, Sijia, et al.
Published: (2026)
by: Luo, Sijia, et al.
Published: (2026)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
by: Wu, Da, et al.
Published: (2023)
by: Wu, Da, et al.
Published: (2023)
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
by: Lin, Huawei, et al.
Published: (2026)
by: Lin, Huawei, et al.
Published: (2026)
Dynamical Systems Theory Behind a Hierarchical Reasoning Model
by: Es'kin, Vasiliy A., et al.
Published: (2026)
by: Es'kin, Vasiliy A., et al.
Published: (2026)
The Recurrent Sticky Hierarchical Dirichlet Process Hidden Markov Model
by: Słupiński, Mikołaj, et al.
Published: (2024)
by: Słupiński, Mikołaj, et al.
Published: (2024)
Large-Step Training Dynamics of a Two-Factor Linear Transformer Model
by: Balasubramanian, Krishnakumar
Published: (2026)
by: Balasubramanian, Krishnakumar
Published: (2026)
Planning Neural Dynamics with Lie Group Embedding through Supervised Projective Manifold Learning
by: Wang, Tianwei, et al.
Published: (2026)
by: Wang, Tianwei, et al.
Published: (2026)
Position: Why a Dynamical Systems Perspective is Needed to Advance Time Series Modeling
by: Durstewitz, Daniel, et al.
Published: (2026)
by: Durstewitz, Daniel, et al.
Published: (2026)
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
by: Kiselev, O. M.
Published: (2026)
by: Kiselev, O. M.
Published: (2026)
Deep learning and the rate of approximation by flows
by: Cheng, Jingpu, et al.
Published: (2026)
by: Cheng, Jingpu, et al.
Published: (2026)
Language models as master equation solvers
by: Liu, Chuanbo, et al.
Published: (2023)
by: Liu, Chuanbo, et al.
Published: (2023)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
by: Somayajula, Sai Ashish, et al.
Published: (2024)
by: Somayajula, Sai Ashish, et al.
Published: (2024)
Stable Anisotropic Regularization
by: Rudman, William, et al.
Published: (2023)
by: Rudman, William, et al.
Published: (2023)
A Hybrid Approach of Transfer Learning and Physics-Informed Modeling: Improving Dissolved Oxygen Concentration Prediction in an Industrial Wastewater Treatment Plant
by: Koksal, Ece S., et al.
Published: (2024)
by: Koksal, Ece S., et al.
Published: (2024)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
MemMamba: Rethinking Memory Patterns in State Space Model
by: Wang, Youjin, et al.
Published: (2025)
by: Wang, Youjin, et al.
Published: (2025)
Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling
by: Shapira, Eilam, et al.
Published: (2026)
by: Shapira, Eilam, et al.
Published: (2026)
Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models
by: Kempton, Tom, et al.
Published: (2025)
by: Kempton, Tom, et al.
Published: (2025)
On the Limitations of Fractal Dimension as a Measure of Generalization
by: Tan, Charlie B., et al.
Published: (2024)
by: Tan, Charlie B., et al.
Published: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Uncovering the Computational Roles of Nonlinearity in Sequence Modeling Using Almost-Linear RNNs
by: Brenner, Manuel, et al.
Published: (2025)
by: Brenner, Manuel, et al.
Published: (2025)
Mnemosyne: An Unsupervised, Human-Inspired Long-Term Memory Architecture for Edge-Based LLMs
by: Jonelagadda, Aneesh, et al.
Published: (2025)
by: Jonelagadda, Aneesh, et al.
Published: (2025)
FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast
by: Bogdanov, Igor, et al.
Published: (2026)
by: Bogdanov, Igor, et al.
Published: (2026)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
by: Takahashi, Hiroshi, et al.
Published: (2026)
by: Takahashi, Hiroshi, et al.
Published: (2026)
Learning phase-space flows using time-discrete implicit Runge-Kutta PINNs
by: Corral, Álvaro Fernández, et al.
Published: (2024)
by: Corral, Álvaro Fernández, et al.
Published: (2024)
Horizon-Constrained Rashomon Sets for Chaotic Forecasting
by: Kale, Gauri, et al.
Published: (2026)
by: Kale, Gauri, et al.
Published: (2026)
Similar Items
-
LongSSM: On the Length Extension of State-space Models in Language Modelling
by: Wang, Shida
Published: (2024) -
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
by: Wang, Shida, et al.
Published: (2023) -
Hybrid Energy-Based Models for Physical AI: Provably Stable Identification of Port-Hamiltonian Dynamics
by: Betteti, Simone, et al.
Published: (2026) -
Learning Macroscopic Dynamics from Partial Microscopic Observations
by: Chen, Mengyi, et al.
Published: (2024) -
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)