SMR: State Memory Replay for Long Sequence Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qi, Biqing, Gao, Junqi, Zhang, Kaiyan, Li, Dong, Liu, Jianxing, Wu, Ligang, Zhou, Bowen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929378242330624
author Qi, Biqing
Gao, Junqi
Zhang, Kaiyan
Li, Dong
Liu, Jianxing
Wu, Ligang
Zhou, Bowen
author_facet Qi, Biqing
Gao, Junqi
Zhang, Kaiyan
Li, Dong
Liu, Jianxing
Wu, Ligang
Zhou, Bowen
contents Despite the promising performance of state space models (SSMs) in long sequence modeling, limitations still exist. Advanced SSMs like S5 and S6 (Mamba) in addressing non-uniform sampling, their recursive structures impede efficient SSM computation via convolution. To overcome compatibility limitations in parallel convolutional computation, this paper proposes a novel non-recursive non-uniform sample processing strategy. Theoretical analysis of SSMs through the lens of Event-Triggered Control (ETC) theory reveals the Non-Stable State (NSS) problem, where deviations from sampling point requirements lead to error transmission and accumulation, causing the divergence of the SSM's hidden state. Our analysis further reveals that adjustments of input sequences with early memories can mitigate the NSS problem, achieving Sampling Step Adaptation (SSA). Building on this insight, we introduce a simple yet effective plug-and-play mechanism, State Memory Replay (SMR), which utilizes learnable memories to adjust the current state with multi-step information for generalization at sampling points different from those in the training data. This enables SSMs to stably model varying sampling points. Experiments on long-range modeling tasks in autoregressive language modeling and Long Range Arena demonstrate the general effectiveness of the SMR mechanism for a series of SSM models.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17534
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SMR: State Memory Replay for Long Sequence Modeling
Qi, Biqing
Gao, Junqi
Zhang, Kaiyan
Li, Dong
Liu, Jianxing
Wu, Ligang
Zhou, Bowen
Machine Learning
Despite the promising performance of state space models (SSMs) in long sequence modeling, limitations still exist. Advanced SSMs like S5 and S6 (Mamba) in addressing non-uniform sampling, their recursive structures impede efficient SSM computation via convolution. To overcome compatibility limitations in parallel convolutional computation, this paper proposes a novel non-recursive non-uniform sample processing strategy. Theoretical analysis of SSMs through the lens of Event-Triggered Control (ETC) theory reveals the Non-Stable State (NSS) problem, where deviations from sampling point requirements lead to error transmission and accumulation, causing the divergence of the SSM's hidden state. Our analysis further reveals that adjustments of input sequences with early memories can mitigate the NSS problem, achieving Sampling Step Adaptation (SSA). Building on this insight, we introduce a simple yet effective plug-and-play mechanism, State Memory Replay (SMR), which utilizes learnable memories to adjust the current state with multi-step information for generalization at sampling points different from those in the training data. This enables SSMs to stably model varying sampling points. Experiments on long-range modeling tasks in autoregressive language modeling and Long Range Arena demonstrate the general effectiveness of the SMR mechanism for a series of SSM models.
title SMR: State Memory Replay for Long Sequence Modeling
topic Machine Learning
url https://arxiv.org/abs/2405.17534