The Hidden Attention of Mamba Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917626814398464 |
|---|---|
| author | Ali, Ameen Zimerman, Itamar Wolf, Lior |
| author_facet | Ali, Ameen Zimerman, Itamar Wolf, Lior |
| contents | The Mamba layer offers an efficient selective state space model (SSM) that is highly effective in modeling multiple domains, including NLP, long-range sequence processing, and computer vision. Selective SSMs are viewed as dual models, in which one trains in parallel on the entire sequence via an IO-aware parallel scan, and deploys in an autoregressive manner. We add a third view and show that such models can be viewed as attention-driven models. This new perspective enables us to empirically and theoretically compare the underlying mechanisms to that of the self-attention layers in transformers and allows us to peer inside the inner workings of the Mamba model with explainability methods. Our code is publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_01590 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | The Hidden Attention of Mamba Models Ali, Ameen Zimerman, Itamar Wolf, Lior Machine Learning F.2.2, I.2.7 F.2.2; I.2.7 The Mamba layer offers an efficient selective state space model (SSM) that is highly effective in modeling multiple domains, including NLP, long-range sequence processing, and computer vision. Selective SSMs are viewed as dual models, in which one trains in parallel on the entire sequence via an IO-aware parallel scan, and deploys in an autoregressive manner. We add a third view and show that such models can be viewed as attention-driven models. This new perspective enables us to empirically and theoretically compare the underlying mechanisms to that of the self-attention layers in transformers and allows us to peer inside the inner workings of the Mamba model with explainability methods. Our code is publicly available. |
| title | The Hidden Attention of Mamba Models |
| topic | Machine Learning F.2.2, I.2.7 F.2.2; I.2.7 |
| url | https://arxiv.org/abs/2403.01590 |