Mamba-R: Vision Mamba ALSO Needs Registers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Feng, Wang, Jiahao, Ren, Sucheng, Wei, Guoyizhe, Mei, Jieru, Shao, Wei, Zhou, Yuyin, Yuille, Alan, Xie, Cihang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913866265395200
author Wang, Feng
Wang, Jiahao
Ren, Sucheng
Wei, Guoyizhe
Mei, Jieru
Shao, Wei
Zhou, Yuyin
Yuille, Alan
Xie, Cihang
author_facet Wang, Feng
Wang, Jiahao
Ren, Sucheng
Wei, Guoyizhe
Mei, Jieru
Shao, Wei
Zhou, Yuyin
Yuille, Alan
Xie, Cihang
contents Similar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in low-information background areas of images, appear much more severe in Vision Mamba -- they exist prevalently even with the tiny-sized model and activate extensively across background regions. To mitigate this issue, we follow the prior solution of introducing register tokens into Vision Mamba. To better cope with Mamba blocks' uni-directional inference paradigm, two key modifications are introduced: 1) evenly inserting registers throughout the input token sequence, and 2) recycling registers for final decision predictions. We term this new architecture Mamba-R. Qualitative observations suggest, compared to vanilla Vision Mamba, Mamba-R's feature maps appear cleaner and more focused on semantically meaningful regions. Quantitatively, Mamba-R attains stronger performance and scales better. For example, on the ImageNet benchmark, our base-size Mamba-R attains 83.0% accuracy, significantly outperforming Vim-B's 81.8%; furthermore, we provide the first successful scaling to the large model size (i.e., with 341M parameters), attaining a competitive accuracy of 83.6% (84.5% if finetuned with 384x384 inputs). Additional validation on the downstream semantic segmentation task also supports Mamba-R's efficacy. Code is available at https://github.com/wangf3014/Mamba-Reg.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14858
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mamba-R: Vision Mamba ALSO Needs Registers
Wang, Feng
Wang, Jiahao
Ren, Sucheng
Wei, Guoyizhe
Mei, Jieru
Shao, Wei
Zhou, Yuyin
Yuille, Alan
Xie, Cihang
Computer Vision and Pattern Recognition
Similar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in low-information background areas of images, appear much more severe in Vision Mamba -- they exist prevalently even with the tiny-sized model and activate extensively across background regions. To mitigate this issue, we follow the prior solution of introducing register tokens into Vision Mamba. To better cope with Mamba blocks' uni-directional inference paradigm, two key modifications are introduced: 1) evenly inserting registers throughout the input token sequence, and 2) recycling registers for final decision predictions. We term this new architecture Mamba-R. Qualitative observations suggest, compared to vanilla Vision Mamba, Mamba-R's feature maps appear cleaner and more focused on semantically meaningful regions. Quantitatively, Mamba-R attains stronger performance and scales better. For example, on the ImageNet benchmark, our base-size Mamba-R attains 83.0% accuracy, significantly outperforming Vim-B's 81.8%; furthermore, we provide the first successful scaling to the large model size (i.e., with 341M parameters), attaining a competitive accuracy of 83.6% (84.5% if finetuned with 384x384 inputs). Additional validation on the downstream semantic segmentation task also supports Mamba-R's efficacy. Code is available at https://github.com/wangf3014/Mamba-Reg.
title Mamba-R: Vision Mamba ALSO Needs Registers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.14858