Memory Determines Learning Direction: A Theory of Gradient-Based Optimization in State Space Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guan, JingChuan, Kubota, Tomoyuki, Kuniyoshi, Yasuo, Nakajima, Kohei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How noise affects memory in linear recurrent networks
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
Training Spiking Neural Networks via Augmented Direct Feedback Alignment
von: Zhang, Yongbo, et al.
Veröffentlicht: (2024)
von: Zhang, Yongbo, et al.
Veröffentlicht: (2024)
Transformation Categorization Based on Group Decomposition Theory Using Parameter Division
von: Komatsu, Takayuki, et al.
Veröffentlicht: (2026)
von: Komatsu, Takayuki, et al.
Veröffentlicht: (2026)
Informational Embodiment: Computational role of information structure in codes and robots
von: Pitti, Alexandre, et al.
Veröffentlicht: (2024)
von: Pitti, Alexandre, et al.
Veröffentlicht: (2024)
Topology Matters: A Cautionary Case Study of Graph SSL on Neuro-Inspired Benchmarks
von: Carlon, May Kristine Jonson, et al.
Veröffentlicht: (2026)
von: Carlon, May Kristine Jonson, et al.
Veröffentlicht: (2026)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
Backward-Friendly Optimization: Training Large Language Models with Approximate Gradients under Memory Constraints
von: Yang, Jing, et al.
Veröffentlicht: (2025)
von: Yang, Jing, et al.
Veröffentlicht: (2025)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
von: Guan, Lei, et al.
Veröffentlicht: (2023)
von: Guan, Lei, et al.
Veröffentlicht: (2023)
Multifunctional physical reservoir computing in soft tensegrity robots
von: Terajima, Ryo, et al.
Veröffentlicht: (2025)
von: Terajima, Ryo, et al.
Veröffentlicht: (2025)
State Space Models over Directed Graphs
von: She, Junzhi, et al.
Veröffentlicht: (2025)
von: She, Junzhi, et al.
Veröffentlicht: (2025)
Learning Associative Memories with Gradient Descent
von: Cabannes, Vivien, et al.
Veröffentlicht: (2024)
von: Cabannes, Vivien, et al.
Veröffentlicht: (2024)
Designing Chaotic Attractors: A Semi-supervised Approach
von: Kabayama, Tempei, et al.
Veröffentlicht: (2024)
von: Kabayama, Tempei, et al.
Veröffentlicht: (2024)
SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting
von: Qiu, Haomiao, et al.
Veröffentlicht: (2025)
von: Qiu, Haomiao, et al.
Veröffentlicht: (2025)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
von: Xie, Shifeng, et al.
Veröffentlicht: (2025)
von: Xie, Shifeng, et al.
Veröffentlicht: (2025)
ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning
von: Kagaya, Tomoyuki, et al.
Veröffentlicht: (2025)
von: Kagaya, Tomoyuki, et al.
Veröffentlicht: (2025)
Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization
von: Shen, Qianli, et al.
Veröffentlicht: (2024)
von: Shen, Qianli, et al.
Veröffentlicht: (2024)
Reservoir Computing Generalized
von: Kubota, Tomoyuki, et al.
Veröffentlicht: (2024)
von: Kubota, Tomoyuki, et al.
Veröffentlicht: (2024)
Gradient Extrapolation-Based Policy Optimization
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
Can Transformers Predict Vibrations?
von: Kuniyoshi, Fusataka, et al.
Veröffentlicht: (2024)
von: Kuniyoshi, Fusataka, et al.
Veröffentlicht: (2024)
Exploiting Chaotic Dynamics as Deep Neural Networks
von: Liu, Shuhong, et al.
Veröffentlicht: (2024)
von: Liu, Shuhong, et al.
Veröffentlicht: (2024)
Offline Model-Based Optimization via Policy-Guided Gradient Search
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
Gradient-Optimized Fuzzy Classifier: A Benchmark Study Against State-of-the-Art Models
von: Sieverding, Magnus, et al.
Veröffentlicht: (2025)
von: Sieverding, Magnus, et al.
Veröffentlicht: (2025)
Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories
von: Xu, Yongzhong
Veröffentlicht: (2026)
von: Xu, Yongzhong
Veröffentlicht: (2026)
MemMamba: Rethinking Memory Patterns in State Space Model
von: Wang, Youjin, et al.
Veröffentlicht: (2025)
von: Wang, Youjin, et al.
Veröffentlicht: (2025)
Mathematical Formalism for Memory Compression in Selective State Space Models
von: Bhat, Siddhanth
Veröffentlicht: (2024)
von: Bhat, Siddhanth
Veröffentlicht: (2024)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
von: Lepel, Olivier, et al.
Veröffentlicht: (2024)
von: Lepel, Olivier, et al.
Veröffentlicht: (2024)
Directed Acyclic Graph Structure Learning from Dynamic Graphs
von: Fan, Shaohua, et al.
Veröffentlicht: (2022)
von: Fan, Shaohua, et al.
Veröffentlicht: (2022)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
von: Galesloot, Maris F. L., et al.
Veröffentlicht: (2025)
von: Galesloot, Maris F. L., et al.
Veröffentlicht: (2025)
Compute-in-Memory Implementation of State Space Models for Event Sequence Processing
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2025)
Reimagining Parameter Space Exploration with Diffusion Models
von: Zhang, Lijun, et al.
Veröffentlicht: (2025)
von: Zhang, Lijun, et al.
Veröffentlicht: (2025)
Permutation-Equivariant 2D State Space Models: Theory and Canonical Architecture for Multivariate Time Series
von: Jeong, Seungwoo, et al.
Veröffentlicht: (2026)
von: Jeong, Seungwoo, et al.
Veröffentlicht: (2026)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
von: Luo, Haocheng, et al.
Veröffentlicht: (2026)
von: Luo, Haocheng, et al.
Veröffentlicht: (2026)
Distributionally Robust Model-based Reinforcement Learning with Large State Spaces
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2023)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2023)
Fine-Tuning Adaptive Stochastic Optimizers: Determining the Optimal Hyperparameter $ε$ via Gradient Magnitude Histogram Analysis
von: Silva, Gustavo, et al.
Veröffentlicht: (2023)
von: Silva, Gustavo, et al.
Veröffentlicht: (2023)
Accelerating LMO-Based Optimization via Implicit Gradient Transport
von: Jang, Won-Jun, et al.
Veröffentlicht: (2026)
von: Jang, Won-Jun, et al.
Veröffentlicht: (2026)
Towards Differentiable Multilevel Optimization: A Gradient-Based Approach
von: Gu, Yuntian, et al.
Veröffentlicht: (2024)
von: Gu, Yuntian, et al.
Veröffentlicht: (2024)
Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling
von: Ceni, Andrea, et al.
Veröffentlicht: (2025)
von: Ceni, Andrea, et al.
Veröffentlicht: (2025)
The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations
von: Lehmann, Matthias
Veröffentlicht: (2024)
von: Lehmann, Matthias
Veröffentlicht: (2024)
Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces
von: Ota, Toshihiro
Veröffentlicht: (2024)
von: Ota, Toshihiro
Veröffentlicht: (2024)
Preference-Based Gradient Estimation for ML-Guided Approximate Combinatorial Optimization
von: Mielke, Arman, et al.
Veröffentlicht: (2025)
von: Mielke, Arman, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How noise affects memory in linear recurrent networks
von: Guan, JingChuan, et al.
Veröffentlicht: (2024) -
Training Spiking Neural Networks via Augmented Direct Feedback Alignment
von: Zhang, Yongbo, et al.
Veröffentlicht: (2024) -
Transformation Categorization Based on Group Decomposition Theory Using Parameter Division
von: Komatsu, Takayuki, et al.
Veröffentlicht: (2026) -
Informational Embodiment: Computational role of information structure in codes and robots
von: Pitti, Alexandre, et al.
Veröffentlicht: (2024) -
Topology Matters: A Cautionary Case Study of Graph SSL on Neuro-Inspired Benchmarks
von: Carlon, May Kristine Jonson, et al.
Veröffentlicht: (2026)