Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yi, Xunpeng, Xu, Han, Zhang, Hao, Tang, Linfeng, Ma, Jiayi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911812033708032
author Yi, Xunpeng
Xu, Han
Zhang, Hao
Tang, Linfeng
Ma, Jiayi
author_facet Yi, Xunpeng
Xu, Han
Zhang, Hao
Tang, Linfeng
Ma, Jiayi
contents Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16387
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
Yi, Xunpeng
Xu, Han
Zhang, Hao
Tang, Linfeng
Ma, Jiayi
Computer Vision and Pattern Recognition
Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF.
title Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.16387