Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909035729518592 |
|---|---|
| author | Wu, Chen Wang, Ling Zheng, Zhuoran Chen, Xiangyu Xia, Jingyuan Jiang, Weidong Zhou, Jiantao |
| author_facet | Wu, Chen Wang, Ling Zheng, Zhuoran Chen, Xiangyu Xia, Jingyuan Jiang, Weidong Zhou, Jiantao |
| contents | Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms with quadratic complexity, hindering the establishment of efficient and semantically interactive joint representations. In this paper, we observe that feature maps from different modalities exhibit semantic-level correlations during feature extraction. This motivates us to develop a more flexible approach enabling dense, semantically-aware deep interactions between modalities. To this end, we propose a novel GDSR framework centered around the Interactive State Space Model. Specifically, we design a cross-modal local scanning mechanism that enables fine-grained semantic interactions between RGB and depth features. Leveraging the Mamba architecture, our framework achieves global modeling with linear complexity. Furthermore, a cross-modal matching transform module is introduced to enhance interactive modeling quality by utilizing representative features from both modalities. Extensive experiments demonstrate competitive performance against state-of-the-art methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11934 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution Wu, Chen Wang, Ling Zheng, Zhuoran Chen, Xiangyu Xia, Jingyuan Jiang, Weidong Zhou, Jiantao Computer Vision and Pattern Recognition Guided depth super-resolution (GDSR) reconstructs HR depth maps from LR inputs with HR RGB guidance. Existing methods either model each modality independently or rely on computationally expensive attention mechanisms with quadratic complexity, hindering the establishment of efficient and semantically interactive joint representations. In this paper, we observe that feature maps from different modalities exhibit semantic-level correlations during feature extraction. This motivates us to develop a more flexible approach enabling dense, semantically-aware deep interactions between modalities. To this end, we propose a novel GDSR framework centered around the Interactive State Space Model. Specifically, we design a cross-modal local scanning mechanism that enables fine-grained semantic interactions between RGB and depth features. Leveraging the Mamba architecture, our framework achieves global modeling with linear complexity. Furthermore, a cross-modal matching transform module is introduced to enhance interactive modeling quality by utilizing representative features from both modalities. Extensive experiments demonstrate competitive performance against state-of-the-art methods. |
| title | Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.11934 |