CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yiming, Ruan, Yuan, Hu, Qinghua, Zhu, Pengfei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909988877762560
author Sun, Yiming
Ruan, Yuan
Hu, Qinghua
Zhu, Pengfei
author_facet Sun, Yiming
Ruan, Yuan
Hu, Qinghua
Zhu, Pengfei
contents Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitly learn rigid semantics through cascaded detection/segmentation models, unable to interactively address diverse semantic target perception needs. We propose CtrlFuse, a controllable image fusion framework that enables interactive dynamic fusion guided by mask prompts. The model integrates a multi-modal feature extractor, a reference prompt encoder (RPE), and a prompt-semantic fusion module (PSFM). The RPE dynamically encodes task-specific semantic prompts by fine-tuning pre-trained segmentation models with input mask guidance, while the PSFM explicitly injects these semantics into fusion features. Through synergistic optimization of parallel segmentation and fusion branches, our method achieves mutual enhancement between task performance and fusion quality. Experiments demonstrate state-of-the-art results in both fusion controllability and segmentation accuracy, with the adapted task branch even outperforming the original segmentation model.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08619
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
Sun, Yiming
Ruan, Yuan
Hu, Qinghua
Zhu, Pengfei
Computer Vision and Pattern Recognition
Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitly learn rigid semantics through cascaded detection/segmentation models, unable to interactively address diverse semantic target perception needs. We propose CtrlFuse, a controllable image fusion framework that enables interactive dynamic fusion guided by mask prompts. The model integrates a multi-modal feature extractor, a reference prompt encoder (RPE), and a prompt-semantic fusion module (PSFM). The RPE dynamically encodes task-specific semantic prompts by fine-tuning pre-trained segmentation models with input mask guidance, while the PSFM explicitly injects these semantics into fusion features. Through synergistic optimization of parallel segmentation and fusion branches, our method achieves mutual enhancement between task performance and fusion quality. Experiments demonstrate state-of-the-art results in both fusion controllability and segmentation accuracy, with the adapted task branch even outperforming the original segmentation model.
title CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.08619