Memory Determines Learning Direction: A Theory of Gradient-Based Optimization in State Space Models
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, JingChuan, Kubota, Tomoyuki, Kuniyoshi, Yasuo, Nakajima, Kohei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How noise affects memory in linear recurrent networks
by: Guan, JingChuan, et al.
Published: (2024)
by: Guan, JingChuan, et al.
Published: (2024)
Training Spiking Neural Networks via Augmented Direct Feedback Alignment
by: Zhang, Yongbo, et al.
Published: (2024)
by: Zhang, Yongbo, et al.
Published: (2024)
Transformation Categorization Based on Group Decomposition Theory Using Parameter Division
by: Komatsu, Takayuki, et al.
Published: (2026)
by: Komatsu, Takayuki, et al.
Published: (2026)
Informational Embodiment: Computational role of information structure in codes and robots
by: Pitti, Alexandre, et al.
Published: (2024)
by: Pitti, Alexandre, et al.
Published: (2024)
Topology Matters: A Cautionary Case Study of Graph SSL on Neuro-Inspired Benchmarks
by: Carlon, May Kristine Jonson, et al.
Published: (2026)
by: Carlon, May Kristine Jonson, et al.
Published: (2026)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026)
by: Terzić, Aleksandar, et al.
Published: (2026)
Backward-Friendly Optimization: Training Large Language Models with Approximate Gradients under Memory Constraints
by: Yang, Jing, et al.
Published: (2025)
by: Yang, Jing, et al.
Published: (2025)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
by: Guan, Lei, et al.
Published: (2023)
by: Guan, Lei, et al.
Published: (2023)
Multifunctional physical reservoir computing in soft tensegrity robots
by: Terajima, Ryo, et al.
Published: (2025)
by: Terajima, Ryo, et al.
Published: (2025)
State Space Models over Directed Graphs
by: She, Junzhi, et al.
Published: (2025)
by: She, Junzhi, et al.
Published: (2025)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Designing Chaotic Attractors: A Semi-supervised Approach
by: Kabayama, Tempei, et al.
Published: (2024)
by: Kabayama, Tempei, et al.
Published: (2024)
SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
by: Qiu, Haomiao, et al.
Published: (2025)
by: Qiu, Haomiao, et al.
Published: (2025)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
by: Xie, Shifeng, et al.
Published: (2025)
by: Xie, Shifeng, et al.
Published: (2025)
ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning
by: Kagaya, Tomoyuki, et al.
Published: (2025)
by: Kagaya, Tomoyuki, et al.
Published: (2025)
Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization
by: Shen, Qianli, et al.
Published: (2024)
by: Shen, Qianli, et al.
Published: (2024)
Reservoir Computing Generalized
by: Kubota, Tomoyuki, et al.
Published: (2024)
by: Kubota, Tomoyuki, et al.
Published: (2024)
Gradient Extrapolation-Based Policy Optimization
by: Swapnil, Ismam Nur, et al.
Published: (2026)
by: Swapnil, Ismam Nur, et al.
Published: (2026)
Can Transformers Predict Vibrations?
by: Kuniyoshi, Fusataka, et al.
Published: (2024)
by: Kuniyoshi, Fusataka, et al.
Published: (2024)
Exploiting Chaotic Dynamics as Deep Neural Networks
by: Liu, Shuhong, et al.
Published: (2024)
by: Liu, Shuhong, et al.
Published: (2024)
Offline Model-Based Optimization via Policy-Guided Gradient Search
by: Chemingui, Yassine, et al.
Published: (2024)
by: Chemingui, Yassine, et al.
Published: (2024)
Gradient-Optimized Fuzzy Classifier: A Benchmark Study Against State-of-the-Art Models
by: Sieverding, Magnus, et al.
Published: (2025)
by: Sieverding, Magnus, et al.
Published: (2025)
Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
MemMamba: Rethinking Memory Patterns in State Space Model
by: Wang, Youjin, et al.
Published: (2025)
by: Wang, Youjin, et al.
Published: (2025)
Mathematical Formalism for Memory Compression in Selective State Space Models
by: Bhat, Siddhanth
Published: (2024)
by: Bhat, Siddhanth
Published: (2024)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
by: Lepel, Olivier, et al.
Published: (2024)
by: Lepel, Olivier, et al.
Published: (2024)
Directed Acyclic Graph Structure Learning from Dynamic Graphs
by: Fan, Shaohua, et al.
Published: (2022)
by: Fan, Shaohua, et al.
Published: (2022)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025)
by: Galesloot, Maris F. L., et al.
Published: (2025)
Compute-in-Memory Implementation of State Space Models for Event Sequence Processing
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Reimagining Parameter Space Exploration with Diffusion Models
by: Zhang, Lijun, et al.
Published: (2025)
by: Zhang, Lijun, et al.
Published: (2025)
Permutation-Equivariant 2D State Space Models: Theory and Canonical Architecture for Multivariate Time Series
by: Jeong, Seungwoo, et al.
Published: (2026)
by: Jeong, Seungwoo, et al.
Published: (2026)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
by: Luo, Haocheng, et al.
Published: (2026)
by: Luo, Haocheng, et al.
Published: (2026)
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
by: Ramesh, Shyam Sundhar, et al.
Published: (2023)
by: Ramesh, Shyam Sundhar, et al.
Published: (2023)
Fine-Tuning Adaptive Stochastic Optimizers: Determining the Optimal Hyperparameter $ε$ via Gradient Magnitude Histogram Analysis
by: Silva, Gustavo, et al.
Published: (2023)
by: Silva, Gustavo, et al.
Published: (2023)
Accelerating LMO-Based Optimization via Implicit Gradient Transport
by: Jang, Won-Jun, et al.
Published: (2026)
by: Jang, Won-Jun, et al.
Published: (2026)
Towards Differentiable Multilevel Optimization: A Gradient-Based Approach
by: Gu, Yuntian, et al.
Published: (2024)
by: Gu, Yuntian, et al.
Published: (2024)
Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling
by: Ceni, Andrea, et al.
Published: (2025)
by: Ceni, Andrea, et al.
Published: (2025)
The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations
by: Lehmann, Matthias
Published: (2024)
by: Lehmann, Matthias
Published: (2024)
Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces
by: Ota, Toshihiro
Published: (2024)
by: Ota, Toshihiro
Published: (2024)
Preference-Based Gradient Estimation for ML-Guided Approximate Combinatorial Optimization
by: Mielke, Arman, et al.
Published: (2025)
by: Mielke, Arman, et al.
Published: (2025)
Similar Items
-
How noise affects memory in linear recurrent networks
by: Guan, JingChuan, et al.
Published: (2024) -
Training Spiking Neural Networks via Augmented Direct Feedback Alignment
by: Zhang, Yongbo, et al.
Published: (2024) -
Transformation Categorization Based on Group Decomposition Theory Using Parameter Division
by: Komatsu, Takayuki, et al.
Published: (2026) -
Informational Embodiment: Computational role of information structure in codes and robots
by: Pitti, Alexandre, et al.
Published: (2024) -
Topology Matters: A Cautionary Case Study of Graph SSL on Neuro-Inspired Benchmarks
by: Carlon, May Kristine Jonson, et al.
Published: (2026)