Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Ninghui, Tosi, Fabio, Wang, Lihui, Han, Jiawei, Bartolomei, Luca, Yao, Zhiting, Poggi, Matteo, Mattoccia, Stefano
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915940737744896
author Xu, Ninghui
Tosi, Fabio
Wang, Lihui
Han, Jiawei
Bartolomei, Luca
Yao, Zhiting
Poggi, Matteo
Mattoccia, Stefano
author_facet Xu, Ninghui
Tosi, Fabio
Wang, Lihui
Han, Jiawei
Bartolomei, Luca
Yao, Zhiting
Poggi, Matteo
Mattoccia, Stefano
contents Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternative visual representation with higher dynamic range free from such limitations. The complementary characteristics of the two modalities make event-frame asymmetric stereo promising for reliable 3D perception under fast motion and challenging illumination. However, the modality gap often leads to marginalization of domain-specific cues essential for cross-modal stereo matching. In this paper, we introduce Bi-CMPStereo, a novel bidirectional cross-modal prompting framework that fully exploits semantic and structural features from both domains for robust matching. Our approach learns finely aligned stereo representations within a target canonical space and integrates complementary representations by projecting each modality into both event and frame domains. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods in accuracy and generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15312
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
Xu, Ninghui
Tosi, Fabio
Wang, Lihui
Han, Jiawei
Bartolomei, Luca
Yao, Zhiting
Poggi, Matteo
Mattoccia, Stefano
Computer Vision and Pattern Recognition
Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternative visual representation with higher dynamic range free from such limitations. The complementary characteristics of the two modalities make event-frame asymmetric stereo promising for reliable 3D perception under fast motion and challenging illumination. However, the modality gap often leads to marginalization of domain-specific cues essential for cross-modal stereo matching. In this paper, we introduce Bi-CMPStereo, a novel bidirectional cross-modal prompting framework that fully exploits semantic and structural features from both domains for robust matching. Our approach learns finely aligned stereo representations within a target canonical space and integrates complementary representations by projecting each modality into both event and frame domains. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods in accuracy and generalization.
title Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.15312