ReMamber: Referring Image Segmentation with Mamba Twister

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Yuhuan, Ma, Chaofan, Yao, Jiangchao, Zhong, Zhun, Zhang, Ya, Wang, Yanfeng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911965934256128
author Yang, Yuhuan
Ma, Chaofan
Yao, Jiangchao
Zhong, Zhun
Zhang, Ya
Wang, Yanfeng
author_facet Yang, Yuhuan
Ma, Chaofan
Yao, Jiangchao
Zhong, Zhun
Zhang, Ya
Wang, Yanfeng
contents Referring Image Segmentation~(RIS) leveraging transformers has achieved great success on the interpretation of complex visual-language tasks. However, the quadratic computation cost makes it resource-consuming in capturing long-range visual-language dependencies. Fortunately, Mamba addresses this with efficient linear complexity in processing. However, directly applying Mamba to multi-modal interactions presents challenges, primarily due to inadequate channel interactions for the effective fusion of multi-modal data. In this paper, we propose ReMamber, a novel RIS architecture that integrates the power of Mamba with a multi-modal Mamba Twister block. The Mamba Twister explicitly models image-text interaction, and fuses textual and visual features through its unique channel and spatial twisting mechanism. We achieve competitive results on three challenging benchmarks with a simple and efficient architecture. Moreover, we conduct thorough analyses of ReMamber and discuss other fusion designs using Mamba. These provide valuable perspectives for future research. The code has been released at: https://github.com/yyh-rain-song/ReMamber.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ReMamber: Referring Image Segmentation with Mamba Twister
Yang, Yuhuan
Ma, Chaofan
Yao, Jiangchao
Zhong, Zhun
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Referring Image Segmentation~(RIS) leveraging transformers has achieved great success on the interpretation of complex visual-language tasks. However, the quadratic computation cost makes it resource-consuming in capturing long-range visual-language dependencies. Fortunately, Mamba addresses this with efficient linear complexity in processing. However, directly applying Mamba to multi-modal interactions presents challenges, primarily due to inadequate channel interactions for the effective fusion of multi-modal data. In this paper, we propose ReMamber, a novel RIS architecture that integrates the power of Mamba with a multi-modal Mamba Twister block. The Mamba Twister explicitly models image-text interaction, and fuses textual and visual features through its unique channel and spatial twisting mechanism. We achieve competitive results on three challenging benchmarks with a simple and efficient architecture. Moreover, we conduct thorough analyses of ReMamber and discuss other fusion designs using Mamba. These provide valuable perspectives for future research. The code has been released at: https://github.com/yyh-rain-song/ReMamber.
title ReMamber: Referring Image Segmentation with Mamba Twister
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2403.17839