Task-driven Image Fusion with Learnable Fusion Loss

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Haowen, Zhang, Jiangshe, Zhao, Zixiang, Wu, Yichen, Deng, Lilun, Cui, Yukun, Feng, Tao, Xu, Shuang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909548789366784
author Bai, Haowen
Zhang, Jiangshe
Zhao, Zixiang
Wu, Yichen
Deng, Lilun
Cui, Yukun
Feng, Tao
Xu, Shuang
author_facet Bai, Haowen
Zhang, Jiangshe
Zhao, Zixiang
Wu, Yichen
Deng, Lilun
Cui, Yukun
Feng, Tao
Xu, Shuang
contents Multi-modal image fusion aggregates information from multiple sensor sources, achieving superior visual quality and perceptual features compared to single-source images, often improving downstream tasks. However, current fusion methods for downstream tasks still use predefined fusion objectives that potentially mismatch the downstream tasks, limiting adaptive guidance and reducing model flexibility. To address this, we propose Task-driven Image Fusion (TDFusion), a fusion framework incorporating a learnable fusion loss guided by task loss. Specifically, our fusion loss includes learnable parameters modeled by a neural network called the loss generation module. This module is supervised by the downstream task loss in a meta-learning manner. The learning objective is to minimize the task loss of fused images after optimizing the fusion module with the fusion loss. Iterative updates between the fusion module and the loss module ensure that the fusion network evolves toward minimizing task loss, guiding the fusion process toward the task objectives. TDFusion's training relies entirely on the downstream task loss, making it adaptable to any specific task. It can be applied to any architecture of fusion and task networks. Experiments demonstrate TDFusion's performance through fusion experiments conducted on four different datasets, in addition to evaluations on semantic segmentation and object detection tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03240
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Task-driven Image Fusion with Learnable Fusion Loss
Bai, Haowen
Zhang, Jiangshe
Zhao, Zixiang
Wu, Yichen
Deng, Lilun
Cui, Yukun
Feng, Tao
Xu, Shuang
Computer Vision and Pattern Recognition
Multi-modal image fusion aggregates information from multiple sensor sources, achieving superior visual quality and perceptual features compared to single-source images, often improving downstream tasks. However, current fusion methods for downstream tasks still use predefined fusion objectives that potentially mismatch the downstream tasks, limiting adaptive guidance and reducing model flexibility. To address this, we propose Task-driven Image Fusion (TDFusion), a fusion framework incorporating a learnable fusion loss guided by task loss. Specifically, our fusion loss includes learnable parameters modeled by a neural network called the loss generation module. This module is supervised by the downstream task loss in a meta-learning manner. The learning objective is to minimize the task loss of fused images after optimizing the fusion module with the fusion loss. Iterative updates between the fusion module and the loss module ensure that the fusion network evolves toward minimizing task loss, guiding the fusion process toward the task objectives. TDFusion's training relies entirely on the downstream task loss, making it adaptable to any specific task. It can be applied to any architecture of fusion and task networks. Experiments demonstrate TDFusion's performance through fusion experiments conducted on four different datasets, in addition to evaluations on semantic segmentation and object detection tasks.
title Task-driven Image Fusion with Learnable Fusion Loss
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.03240