ReMamba: Equip Mamba with Effective Long-Sequence Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910768017965056 |
|---|---|
| author | Yuan, Danlong Liu, Jiahao Li, Bei Zhang, Huishuai Wang, Jingang Cai, Xunliang Zhao, Dongyan |
| author_facet | Yuan, Danlong Liu, Jiahao Li, Bei Zhang, Huishuai Wang, Jingang Cai, Xunliang Zhao, Dongyan |
| contents | While the Mamba architecture demonstrates superior inference efficiency and competitive performance on short-context natural language processing (NLP) tasks, empirical evidence suggests its capacity to comprehend long contexts is limited compared to transformer-based models. In this study, we investigate the long-context efficiency issues of the Mamba models and propose ReMamba, which enhances Mamba's ability to comprehend long contexts. ReMamba incorporates selective compression and adaptation techniques within a two-stage re-forward process, incurring minimal additional inference costs overhead. Experimental results on the LongBench and L-Eval benchmarks demonstrate ReMamba's efficacy, improving over the baselines by 3.2 and 1.6 points, respectively, and attaining performance almost on par with same-size transformer models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_15496 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ReMamba: Equip Mamba with Effective Long-Sequence Modeling Yuan, Danlong Liu, Jiahao Li, Bei Zhang, Huishuai Wang, Jingang Cai, Xunliang Zhao, Dongyan Computation and Language While the Mamba architecture demonstrates superior inference efficiency and competitive performance on short-context natural language processing (NLP) tasks, empirical evidence suggests its capacity to comprehend long contexts is limited compared to transformer-based models. In this study, we investigate the long-context efficiency issues of the Mamba models and propose ReMamba, which enhances Mamba's ability to comprehend long contexts. ReMamba incorporates selective compression and adaptation techniques within a two-stage re-forward process, incurring minimal additional inference costs overhead. Experimental results on the LongBench and L-Eval benchmarks demonstrate ReMamba's efficacy, improving over the baselines by 3.2 and 1.6 points, respectively, and attaining performance almost on par with same-size transformer models. |
| title | ReMamba: Equip Mamba with Effective Long-Sequence Modeling |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2408.15496 |