CFMW: Cross-modality Fusion Mamba for Robust Object Detection under Adverse Weather

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Haoyuan, Hu, Qi, Zhou, Binjia, Yao, You, Lin, Jiacheng, Yang, Kailun, Chen, Peng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909677834469376
author Li, Haoyuan
Hu, Qi
Zhou, Binjia
Yao, You
Lin, Jiacheng
Yang, Kailun
Chen, Peng
author_facet Li, Haoyuan
Hu, Qi
Zhou, Binjia
Yao, You
Lin, Jiacheng
Yang, Kailun
Chen, Peng
contents Visible-infrared image pairs provide complementary information, enhancing the reliability and robustness of object detection applications in real-world scenarios. However, most existing methods face challenges in maintaining robustness under complex weather conditions, which limits their applicability. Meanwhile, the reliance on attention mechanisms in modality fusion introduces significant computational complexity and storage overhead, particularly when dealing with high-resolution images. To address these challenges, we propose the Cross-modality Fusion Mamba with Weather-removal (CFMW) to augment stability and cost-effectiveness under adverse weather conditions. Leveraging the proposed Perturbation-Adaptive Diffusion Model (PADM) and Cross-modality Fusion Mamba (CFM) modules, CFMW is able to reconstruct visual features affected by adverse weather, enriching the representation of image details. With efficient architecture design, CFMW is 3 times faster than Transformer-style fusion (e.g., CFT). To bridge the gap in relevant datasets, we construct a new Severe Weather Visible-Infrared (SWVI) dataset, encompassing diverse adverse weather scenarios such as rain, haze, and snow. The dataset contains 64,281 paired visible-infrared images, providing a valuable resource for future research. Extensive experiments on public datasets (i.e., M3FD and LLVIP) and the newly constructed SWVI dataset conclusively demonstrate that CFMW achieves state-of-the-art detection performance. Both the dataset and source code will be made publicly available at https://github.com/lhy-zjut/CFMW.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16302
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CFMW: Cross-modality Fusion Mamba for Robust Object Detection under Adverse Weather
Li, Haoyuan
Hu, Qi
Zhou, Binjia
Yao, You
Lin, Jiacheng
Yang, Kailun
Chen, Peng
Computer Vision and Pattern Recognition
Multimedia
Robotics
Image and Video Processing
Visible-infrared image pairs provide complementary information, enhancing the reliability and robustness of object detection applications in real-world scenarios. However, most existing methods face challenges in maintaining robustness under complex weather conditions, which limits their applicability. Meanwhile, the reliance on attention mechanisms in modality fusion introduces significant computational complexity and storage overhead, particularly when dealing with high-resolution images. To address these challenges, we propose the Cross-modality Fusion Mamba with Weather-removal (CFMW) to augment stability and cost-effectiveness under adverse weather conditions. Leveraging the proposed Perturbation-Adaptive Diffusion Model (PADM) and Cross-modality Fusion Mamba (CFM) modules, CFMW is able to reconstruct visual features affected by adverse weather, enriching the representation of image details. With efficient architecture design, CFMW is 3 times faster than Transformer-style fusion (e.g., CFT). To bridge the gap in relevant datasets, we construct a new Severe Weather Visible-Infrared (SWVI) dataset, encompassing diverse adverse weather scenarios such as rain, haze, and snow. The dataset contains 64,281 paired visible-infrared images, providing a valuable resource for future research. Extensive experiments on public datasets (i.e., M3FD and LLVIP) and the newly constructed SWVI dataset conclusively demonstrate that CFMW achieves state-of-the-art detection performance. Both the dataset and source code will be made publicly available at https://github.com/lhy-zjut/CFMW.
title CFMW: Cross-modality Fusion Mamba for Robust Object Detection under Adverse Weather
topic Computer Vision and Pattern Recognition
Multimedia
Robotics
Image and Video Processing
url https://arxiv.org/abs/2404.16302