MambaIRv2: Attentive State Space Restoration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Hang, Guo, Yong, Zha, Yaohua, Zhang, Yulun, Li, Wenbo, Dai, Tao, Xia, Shu-Tao, Li, Yawei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929752148803584
author Guo, Hang
Guo, Yong
Zha, Yaohua
Zhang, Yulun
Li, Wenbo
Dai, Tao
Xia, Shu-Tao
Li, Yawei
author_facet Guo, Hang
Guo, Yong
Zha, Yaohua
Zhang, Yulun
Li, Wenbo
Dai, Tao
Xia, Shu-Tao
Li, Yawei
contents The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts the full utilization of pixels across the image and thus presents new challenges in image restoration. In this work, we propose MambaIRv2, which equips Mamba with the non-causal modeling ability similar to ViTs to reach the attentive state space restoration model. Specifically, the proposed attentive state-space equation allows to attend beyond the scanned sequence and facilitate image unfolding with just one single scan. Moreover, we further introduce a semantic-guided neighboring mechanism to encourage interaction between distant but similar pixels. Extensive experiments show our MambaIRv2 outperforms SRFormer by even 0.35dB PSNR for lightweight SR even with 9.3\% less parameters and suppresses HAT on classic SR by up to 0.29dB. Code is available at https://github.com/csguoh/MambaIR.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15269
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MambaIRv2: Attentive State Space Restoration
Guo, Hang
Guo, Yong
Zha, Yaohua
Zhang, Yulun
Li, Wenbo
Dai, Tao
Xia, Shu-Tao
Li, Yawei
Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts the full utilization of pixels across the image and thus presents new challenges in image restoration. In this work, we propose MambaIRv2, which equips Mamba with the non-causal modeling ability similar to ViTs to reach the attentive state space restoration model. Specifically, the proposed attentive state-space equation allows to attend beyond the scanned sequence and facilitate image unfolding with just one single scan. Moreover, we further introduce a semantic-guided neighboring mechanism to encourage interaction between distant but similar pixels. Extensive experiments show our MambaIRv2 outperforms SRFormer by even 0.35dB PSNR for lightweight SR even with 9.3\% less parameters and suppresses HAT on classic SR by up to 0.29dB. Code is available at https://github.com/csguoh/MambaIR.
title MambaIRv2: Attentive State Space Restoration
topic Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.15269