Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Tianyao, Xiang, Dawei, Ding, Tianqi, Fang, Xiang, Qi, Yijiashun, Zhao, Zunduo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908538912112640
author Sun, Tianyao
Xiang, Dawei
Ding, Tianqi
Fang, Xiang
Qi, Yijiashun
Zhao, Zunduo
author_facet Sun, Tianyao
Xiang, Dawei
Ding, Tianqi
Fang, Xiang
Qi, Yijiashun
Zhao, Zunduo
contents Infrared and visible image fusion (IVIF) is a fundamental task in multi-modal perception that aims to integrate complementary structural and textural cues from different spectral domains. In this paper, we propose FusionNet, a novel end-to-end fusion framework that explicitly models inter-modality interaction and enhances task-critical regions. FusionNet introduces a modality-aware attention mechanism that dynamically adjusts the contribution of infrared and visible features based on their discriminative capacity. To achieve fine-grained, interpretable fusion, we further incorporate a pixel-wise alpha blending module, which learns spatially-varying fusion weights in an adaptive and content-aware manner. Moreover, we formulate a target-aware loss that leverages weak ROI supervision to preserve semantic consistency in regions containing important objects (e.g., pedestrians, vehicles). Experiments on the public M3FD dataset demonstrate that FusionNet generates fused images with enhanced semantic preservation, high perceptual quality, and clear interpretability. Our framework provides a general and extensible solution for semantic-aware multi-modal image fusion, with benefits for downstream tasks such as object detection and scene understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision
Sun, Tianyao
Xiang, Dawei
Ding, Tianqi
Fang, Xiang
Qi, Yijiashun
Zhao, Zunduo
Computer Vision and Pattern Recognition
Machine Learning
Infrared and visible image fusion (IVIF) is a fundamental task in multi-modal perception that aims to integrate complementary structural and textural cues from different spectral domains. In this paper, we propose FusionNet, a novel end-to-end fusion framework that explicitly models inter-modality interaction and enhances task-critical regions. FusionNet introduces a modality-aware attention mechanism that dynamically adjusts the contribution of infrared and visible features based on their discriminative capacity. To achieve fine-grained, interpretable fusion, we further incorporate a pixel-wise alpha blending module, which learns spatially-varying fusion weights in an adaptive and content-aware manner. Moreover, we formulate a target-aware loss that leverages weak ROI supervision to preserve semantic consistency in regions containing important objects (e.g., pedestrians, vehicles). Experiments on the public M3FD dataset demonstrate that FusionNet generates fused images with enhanced semantic preservation, high perceptual quality, and clear interpretability. Our framework provides a general and extensible solution for semantic-aware multi-modal image fusion, with benefits for downstream tasks such as object detection and scene understanding.
title Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.11476