Self-supervised Multiplex Consensus Mamba for General Image Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yingying, Zhuang, Rongjin, Zheng, Hui, He, Xuanhua, Cao, Ke, Tu, Xiaotong, Ding, Xinghao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917168739778560
author Wang, Yingying
Zhuang, Rongjin
Zheng, Hui
He, Xuanhua
Cao, Ke
Tu, Xiaotong
Ding, Xinghao
author_facet Wang, Yingying
Zhuang, Rongjin
Zheng, Hui
He, Xuanhua
Cao, Ke
Tu, Xiaotong
Ding, Xinghao
contents Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency-rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-supervised Multiplex Consensus Mamba for General Image Fusion
Wang, Yingying
Zhuang, Rongjin
Zheng, Hui
He, Xuanhua
Cao, Ke
Tu, Xiaotong
Ding, Xinghao
Computer Vision and Pattern Recognition
Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency-rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks.
title Self-supervised Multiplex Consensus Mamba for General Image Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.20921