MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Karami, Mahdi, Behrouz, Ali, Zhong, Peilin, Pascanu, Razvan, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
by: Kacham, Praneeth, et al.
Published: (2023)
by: Kacham, Praneeth, et al.
Published: (2023)
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)
by: Behrouz, Ali, et al.
Published: (2026)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2024)
by: Karami, Mahdi, et al.
Published: (2024)
Efficient Data Selection at Scale via Influence Distillation
by: Nikdan, Mahdi, et al.
Published: (2025)
by: Nikdan, Mahdi, et al.
Published: (2025)
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
Differentially Private Graph Learning via Sensitivity-Bounded Personalized PageRank
by: Epasto, Alessandro, et al.
Published: (2022)
by: Epasto, Alessandro, et al.
Published: (2022)
Graph Mamba: Towards Learning on Graphs with State Space Models
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Auto-Regressive Masked Diffusion Models
by: Karami, Mahdi, et al.
Published: (2026)
by: Karami, Mahdi, et al.
Published: (2026)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Retraining with Predicted Hard Labels Provably Increases Model Accuracy
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
by: Meskin, Asal, et al.
Published: (2025)
by: Meskin, Asal, et al.
Published: (2025)
Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
MambaMixer: Efficient Selective State Space Models with Dual Token and Channel Selection
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
by: Ghodsi, Ali
Published: (2025)
by: Ghodsi, Ali
Published: (2025)
Perturb-and-Project: Differentially Private Similarities and Marginals
by: Cohen-Addad, Vincent, et al.
Published: (2024)
by: Cohen-Addad, Vincent, et al.
Published: (2024)
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
by: Tuo, Kaiwen, et al.
Published: (2025)
by: Tuo, Kaiwen, et al.
Published: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Geometric SSM: LTI State Space Models for Selective Tasks
by: Casti, Umberto, et al.
Published: (2025)
by: Casti, Umberto, et al.
Published: (2025)
Improving the Variance of Differentially Private Randomized Experiments through Clustering
by: Javanmard, Adel, et al.
Published: (2023)
by: Javanmard, Adel, et al.
Published: (2023)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention
by: Solozabal, Ruben, et al.
Published: (2026)
by: Solozabal, Ruben, et al.
Published: (2026)
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Learning Rate Schedules in the Presence of Distribution Shift
by: Fahrbach, Matthew, et al.
Published: (2023)
by: Fahrbach, Matthew, et al.
Published: (2023)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
EEG-SSM: Leveraging State-Space Model for Dementia Detection
by: Tran, Xuan-The, et al.
Published: (2024)
by: Tran, Xuan-The, et al.
Published: (2024)
HiGen: Hierarchical Graph Generative Networks
by: Karami, Mahdi
Published: (2023)
by: Karami, Mahdi
Published: (2023)
Similar Items
-
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025) -
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025) -
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024) -
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
by: Behrouz, Ali, et al.
Published: (2024) -
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)