How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
Fuente:
arXiv
Saved in:
| Main Author: | Ghodsi, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Time-SSM: Simplifying and Unifying State Space Models for Time Series Forecasting
by: Hu, Jiaxi, et al.
Published: (2024)
by: Hu, Jiaxi, et al.
Published: (2024)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026)
by: Terzić, Aleksandar, et al.
Published: (2026)
UniMamba: A Unified Spatial-Temporal Modeling Framework with State-Space and Attention Integration
by: Chen, Xingsheng, et al.
Published: (2026)
by: Chen, Xingsheng, et al.
Published: (2026)
EEG-SSM: Leveraging State-Space Model for Dementia Detection
by: Tran, Xuan-The, et al.
Published: (2024)
by: Tran, Xuan-The, et al.
Published: (2024)
Scalable Graph Self-Supervised Learning
by: Pasand, Ali Saheb, et al.
Published: (2024)
by: Pasand, Ali Saheb, et al.
Published: (2024)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
by: Bick, Aviv, et al.
Published: (2026)
by: Bick, Aviv, et al.
Published: (2026)
LongSSM: On the Length Extension of State-space Models in Language Modelling
by: Wang, Shida
Published: (2024)
by: Wang, Shida
Published: (2024)
FACTS: A Factored State-Space Framework For World Modelling
by: Nanbo, Li, et al.
Published: (2024)
by: Nanbo, Li, et al.
Published: (2024)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
ECGMamba: Towards Efficient ECG Classification with BiSSM
by: Qiang, Yupeng, et al.
Published: (2024)
by: Qiang, Yupeng, et al.
Published: (2024)
Cellular Traffic Prediction via Deep State Space Models with Attention Mechanism
by: Ma, Hui, et al.
Published: (2025)
by: Ma, Hui, et al.
Published: (2025)
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
by: Kong, Jason, et al.
Published: (2026)
by: Kong, Jason, et al.
Published: (2026)
Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Singular Vectors of Attention Heads Align with Features
by: Franco, Gabriel, et al.
Published: (2026)
by: Franco, Gabriel, et al.
Published: (2026)
Sessa: Selective State Space Attention
by: Horbatko, Liubomyr
Published: (2026)
by: Horbatko, Liubomyr
Published: (2026)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Do Attention Heads Compete or Cooperate during Counting?
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
by: Mudarisov, Timur, et al.
Published: (2026)
by: Mudarisov, Timur, et al.
Published: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
A Multi-Modal CNN-LSTM Framework with Multi-Head Attention and Focal Loss for Real-Time Elderly Fall Detection
by: Zhou, Lijie, et al.
Published: (2026)
by: Zhou, Lijie, et al.
Published: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
by: Kim, Wall, et al.
Published: (2026)
by: Kim, Wall, et al.
Published: (2026)
Decision MetaMamba: Enhancing Selective SSM in Offline RL with Heterogeneous Sequence Mixing
by: Kim, Wall, et al.
Published: (2024)
by: Kim, Wall, et al.
Published: (2024)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
by: Otsuka, Hikari, et al.
Published: (2025)
by: Otsuka, Hikari, et al.
Published: (2025)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
by: Bali, Karan, et al.
Published: (2026)
by: Bali, Karan, et al.
Published: (2026)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
by: Sieber, Jerome, et al.
Published: (2024)
by: Sieber, Jerome, et al.
Published: (2024)
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
by: Williams, Jorge L. Ruiz
Published: (2026)
by: Williams, Jorge L. Ruiz
Published: (2026)
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
by: Wang, George, et al.
Published: (2024)
by: Wang, George, et al.
Published: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
by: Park, Sumin, et al.
Published: (2025)
by: Park, Sumin, et al.
Published: (2025)
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
by: Jiang, Xinting, et al.
Published: (2026)
by: Jiang, Xinting, et al.
Published: (2026)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Parkinson's Disease Detection from Resting State EEG using Multi-Head Graph Structure Learning with Gradient Weighted Graph Attention Explanations
by: Neves, Christopher, et al.
Published: (2024)
by: Neves, Christopher, et al.
Published: (2024)
Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
by: Nikiforos, Lorenzo, et al.
Published: (2025)
by: Nikiforos, Lorenzo, et al.
Published: (2025)
OMGPT: A Sequence Modeling Framework for Data-driven Operational Decision Making
by: Wang, Hanzhao, et al.
Published: (2025)
by: Wang, Hanzhao, et al.
Published: (2025)
Chimera: State Space Models Beyond Sequences
by: Lahoti, Aakash, et al.
Published: (2025)
by: Lahoti, Aakash, et al.
Published: (2025)
State Space Models over Directed Graphs
by: She, Junzhi, et al.
Published: (2025)
by: She, Junzhi, et al.
Published: (2025)
Similar Items
-
Time-SSM: Simplifying and Unifying State Space Models for Time Series Forecasting
by: Hu, Jiaxi, et al.
Published: (2024) -
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026) -
UniMamba: A Unified Spatial-Temporal Modeling Framework with State-Space and Attention Integration
by: Chen, Xingsheng, et al.
Published: (2026) -
EEG-SSM: Leveraging State-Space Model for Dementia Detection
by: Tran, Xuan-The, et al.
Published: (2024) -
Scalable Graph Self-Supervised Learning
by: Pasand, Ali Saheb, et al.
Published: (2024)