ReMamba: Equip Mamba with Effective Long-Sequence Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Danlong, Liu, Jiahao, Li, Bei, Zhang, Huishuai, Wang, Jingang, Cai, Xunliang, Zhao, Dongyan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910768017965056
author Yuan, Danlong
Liu, Jiahao
Li, Bei
Zhang, Huishuai
Wang, Jingang
Cai, Xunliang
Zhao, Dongyan
author_facet Yuan, Danlong
Liu, Jiahao
Li, Bei
Zhang, Huishuai
Wang, Jingang
Cai, Xunliang
Zhao, Dongyan
contents While the Mamba architecture demonstrates superior inference efficiency and competitive performance on short-context natural language processing (NLP) tasks, empirical evidence suggests its capacity to comprehend long contexts is limited compared to transformer-based models. In this study, we investigate the long-context efficiency issues of the Mamba models and propose ReMamba, which enhances Mamba's ability to comprehend long contexts. ReMamba incorporates selective compression and adaptation techniques within a two-stage re-forward process, incurring minimal additional inference costs overhead. Experimental results on the LongBench and L-Eval benchmarks demonstrate ReMamba's efficacy, improving over the baselines by 3.2 and 1.6 points, respectively, and attaining performance almost on par with same-size transformer models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15496
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReMamba: Equip Mamba with Effective Long-Sequence Modeling
Yuan, Danlong
Liu, Jiahao
Li, Bei
Zhang, Huishuai
Wang, Jingang
Cai, Xunliang
Zhao, Dongyan
Computation and Language
While the Mamba architecture demonstrates superior inference efficiency and competitive performance on short-context natural language processing (NLP) tasks, empirical evidence suggests its capacity to comprehend long contexts is limited compared to transformer-based models. In this study, we investigate the long-context efficiency issues of the Mamba models and propose ReMamba, which enhances Mamba's ability to comprehend long contexts. ReMamba incorporates selective compression and adaptation techniques within a two-stage re-forward process, incurring minimal additional inference costs overhead. Experimental results on the LongBench and L-Eval benchmarks demonstrate ReMamba's efficacy, improving over the baselines by 3.2 and 1.6 points, respectively, and attaining performance almost on par with same-size transformer models.
title ReMamba: Equip Mamba with Effective Long-Sequence Modeling
topic Computation and Language
url https://arxiv.org/abs/2408.15496