SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Huan, Li, Hui, Xu, Tianyang, Wu, Xiao-Jun, Wang, Rui, Cheng, Chunyang, Kittler, Josef
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918146696282112
author Kang, Huan
Li, Hui
Xu, Tianyang
Wu, Xiao-Jun
Wang, Rui
Cheng, Chunyang
Kittler, Josef
author_facet Kang, Huan
Li, Hui
Xu, Tianyang
Wu, Xiao-Jun
Wang, Rui
Cheng, Chunyang
Kittler, Josef
contents Euclidean representation learning methods have achieved promising results in image fusion tasks, which can be attributed to their clear advantages in handling with linear space. However, data collected from a realistic scene usually has a non-Euclidean structure, evaluating the consistency of latent representations from paired views using Euclidean distance raises challenges. To address this issue, a novel SPD (symmetric positive definite) manifold learning is proposed for multi-modal image fusion, named SMLNet, which extends the image fusion approach from the Euclidean space to the SPD manifolds. Specifically, we encode images according to the Riemannian geometry to exploit their intrinsic statistical correlations, thereby aligning with human visual perception. The SPD matrix fundamentally underpins our network's learning process. Building upon this mathematical foundation, we employ a cross-modal fusion strategy to exploit modality-specific dependencies and augment complementary information. To capture semantic similarity in images' intrinsic space, we further develop an attention module that meticulously processes the cross-modal semantic affinity matrix. Based on this, we design an end-to-end fusion network based on cross-modal manifold learning. Extensive experiments on public datasets demonstrate that our framework exhibits superior performance compared to the current state-of-the-art methods. Our code will be publicly available at https://github.com/Shaoyun2023.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10679
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
Kang, Huan
Li, Hui
Xu, Tianyang
Wu, Xiao-Jun
Wang, Rui
Cheng, Chunyang
Kittler, Josef
Computer Vision and Pattern Recognition
I.4
Euclidean representation learning methods have achieved promising results in image fusion tasks, which can be attributed to their clear advantages in handling with linear space. However, data collected from a realistic scene usually has a non-Euclidean structure, evaluating the consistency of latent representations from paired views using Euclidean distance raises challenges. To address this issue, a novel SPD (symmetric positive definite) manifold learning is proposed for multi-modal image fusion, named SMLNet, which extends the image fusion approach from the Euclidean space to the SPD manifolds. Specifically, we encode images according to the Riemannian geometry to exploit their intrinsic statistical correlations, thereby aligning with human visual perception. The SPD matrix fundamentally underpins our network's learning process. Building upon this mathematical foundation, we employ a cross-modal fusion strategy to exploit modality-specific dependencies and augment complementary information. To capture semantic similarity in images' intrinsic space, we further develop an attention module that meticulously processes the cross-modal semantic affinity matrix. Based on this, we design an end-to-end fusion network based on cross-modal manifold learning. Extensive experiments on public datasets demonstrate that our framework exhibits superior performance compared to the current state-of-the-art methods. Our code will be publicly available at https://github.com/Shaoyun2023.
title SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
topic Computer Vision and Pattern Recognition
I.4
url https://arxiv.org/abs/2411.10679