TeSG: Textual Semantic Guidance for Infrared and Visible Image Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Mingrui, Chen, Xiru, Wei, Xin, Wang, Nannan, Gao, Xinbo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912442311770112
author Zhu, Mingrui
Chen, Xiru
Wei, Xin
Wang, Nannan
Gao, Xinbo
author_facet Zhu, Mingrui
Chen, Xiru
Wei, Xin
Wang, Nannan
Gao, Xinbo
contents Infrared and visible image fusion (IVF) aims to combine complementary information from both image modalities, producing more informative and comprehensive outputs. Recently, text-guided IVF has shown great potential due to its flexibility and versatility. However, the effective integration and utilization of textual semantic information remains insufficiently studied. To tackle these challenges, we introduce textual semantics at two levels: the mask semantic level and the text semantic level, both derived from textual descriptions extracted by large Vision-Language Models (VLMs). Building on this, we propose Textual Semantic Guidance for infrared and visible image fusion, termed TeSG, which guides the image synthesis process in a way that is optimized for downstream tasks such as detection and segmentation. Specifically, TeSG consists of three core components: a Semantic Information Generator (SIG), a Mask-Guided Cross-Attention (MGCA) module, and a Text-Driven Attentional Fusion (TDAF) module. The SIG generates mask and text semantics based on textual descriptions. The MGCA module performs initial attention-based fusion of visual features from both infrared and visible images, guided by mask semantics. Finally, the TDAF module refines the fusion process with gated attention driven by text semantics. Extensive experiments demonstrate the competitiveness of our approach, particularly in terms of performance on downstream tasks, compared to existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16730
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TeSG: Textual Semantic Guidance for Infrared and Visible Image Fusion
Zhu, Mingrui
Chen, Xiru
Wei, Xin
Wang, Nannan
Gao, Xinbo
Computer Vision and Pattern Recognition
Infrared and visible image fusion (IVF) aims to combine complementary information from both image modalities, producing more informative and comprehensive outputs. Recently, text-guided IVF has shown great potential due to its flexibility and versatility. However, the effective integration and utilization of textual semantic information remains insufficiently studied. To tackle these challenges, we introduce textual semantics at two levels: the mask semantic level and the text semantic level, both derived from textual descriptions extracted by large Vision-Language Models (VLMs). Building on this, we propose Textual Semantic Guidance for infrared and visible image fusion, termed TeSG, which guides the image synthesis process in a way that is optimized for downstream tasks such as detection and segmentation. Specifically, TeSG consists of three core components: a Semantic Information Generator (SIG), a Mask-Guided Cross-Attention (MGCA) module, and a Text-Driven Attentional Fusion (TDAF) module. The SIG generates mask and text semantics based on textual descriptions. The MGCA module performs initial attention-based fusion of visual features from both infrared and visible images, guided by mask semantics. Finally, the TDAF module refines the fusion process with gated attention driven by text semantics. Extensive experiments demonstrate the competitiveness of our approach, particularly in terms of performance on downstream tasks, compared to existing state-of-the-art methods.
title TeSG: Textual Semantic Guidance for Infrared and Visible Image Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.16730