Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Chen, Wang, Ling, Zheng, Zhuoran, Chen, Xiangyu, Xia, Jingyuan, Jiang, Weidong, Zhou, Jiantao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909035729518592
author Wu, Chen
Wang, Ling
Zheng, Zhuoran
Chen, Xiangyu
Xia, Jingyuan
Jiang, Weidong
Zhou, Jiantao
author_facet Wu, Chen
Wang, Ling
Zheng, Zhuoran
Chen, Xiangyu
Xia, Jingyuan
Jiang, Weidong
Zhou, Jiantao
contents Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms with quadratic complexity, hindering the establishment of efficient and semantically interactive joint representations. In this paper, we observe that feature maps from different modalities exhibit semantic-level correlations during feature extraction. This motivates us to develop a more flexible approach enabling dense, semantically-aware deep interactions between modalities. To this end, we propose a novel GDSR framework centered around the Interactive State Space Model. Specifically, we design a cross-modal local scanning mechanism that enables fine-grained semantic interactions between RGB and depth features. Leveraging the Mamba architecture, our framework achieves global modeling with linear complexity. Furthermore, a cross-modal matching transform module is introduced to enhance interactive modeling quality by utilizing representative features from both modalities. Extensive experiments demonstrate competitive performance against state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11934
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
Wu, Chen
Wang, Ling
Zheng, Zhuoran
Chen, Xiangyu
Xia, Jingyuan
Jiang, Weidong
Zhou, Jiantao
Computer Vision and Pattern Recognition
Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms with quadratic complexity, hindering the establishment of efficient and semantically interactive joint representations. In this paper, we observe that feature maps from different modalities exhibit semantic-level correlations during feature extraction. This motivates us to develop a more flexible approach enabling dense, semantically-aware deep interactions between modalities. To this end, we propose a novel GDSR framework centered around the Interactive State Space Model. Specifically, we design a cross-modal local scanning mechanism that enables fine-grained semantic interactions between RGB and depth features. Leveraging the Mamba architecture, our framework achieves global modeling with linear complexity. Furthermore, a cross-modal matching transform module is introduced to enhance interactive modeling quality by utilizing representative features from both modalities. Extensive experiments demonstrate competitive performance against state-of-the-art methods.
title Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.11934