GrFormer: A Novel Transformer on Grassmann Manifold for Infrared and Visible Image Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Huan, Li, Hui, Wu, Xiao-Jun, Xu, Tianyang, Wang, Rui, Cheng, Chunyang, Kittler, Josef
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909651473268736
author Kang, Huan
Li, Hui
Wu, Xiao-Jun
Xu, Tianyang
Wang, Rui
Cheng, Chunyang
Kittler, Josef
author_facet Kang, Huan
Li, Hui
Wu, Xiao-Jun
Xu, Tianyang
Wang, Rui
Cheng, Chunyang
Kittler, Josef
contents In the field of image fusion, promising progress has been made by modeling data from different modalities as linear subspaces. However, in practice, the source images are often located in a non-Euclidean space, where the Euclidean methods usually cannot encapsulate the intrinsic topological structure. Typically, the inner product performed in the Euclidean space calculates the algebraic similarity rather than the semantic similarity, which results in undesired attention output and a decrease in fusion performance. While the balance of low-level details and high-level semantics should be considered in infrared and visible image fusion task. To address this issue, in this paper, we propose a novel attention mechanism based on Grassmann manifold for infrared and visible image fusion (GrFormer). Specifically, our method constructs a low-rank subspace mapping through projection constraints on the Grassmann manifold, compressing attention features into subspaces of varying rank levels. This forces the features to decouple into high-frequency details (local low-rank) and low-frequency semantics (global low-rank), thereby achieving multi-scale semantic fusion. Additionally, to effectively integrate the significant information, we develop a cross-modal fusion strategy (CMS) based on a covariance mask to maximise the complementary properties between different modalities and to suppress the features with high correlation, which are deemed redundant. The experimental results demonstrate that our network outperforms SOTA methods both qualitatively and quantitatively on multiple image fusion benchmarks. The codes are available at https://github.com/Shaoyun2023.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GrFormer: A Novel Transformer on Grassmann Manifold for Infrared and Visible Image Fusion
Kang, Huan
Li, Hui
Wu, Xiao-Jun
Xu, Tianyang
Wang, Rui
Cheng, Chunyang
Kittler, Josef
Computer Vision and Pattern Recognition
I.4
In the field of image fusion, promising progress has been made by modeling data from different modalities as linear subspaces. However, in practice, the source images are often located in a non-Euclidean space, where the Euclidean methods usually cannot encapsulate the intrinsic topological structure. Typically, the inner product performed in the Euclidean space calculates the algebraic similarity rather than the semantic similarity, which results in undesired attention output and a decrease in fusion performance. While the balance of low-level details and high-level semantics should be considered in infrared and visible image fusion task. To address this issue, in this paper, we propose a novel attention mechanism based on Grassmann manifold for infrared and visible image fusion (GrFormer). Specifically, our method constructs a low-rank subspace mapping through projection constraints on the Grassmann manifold, compressing attention features into subspaces of varying rank levels. This forces the features to decouple into high-frequency details (local low-rank) and low-frequency semantics (global low-rank), thereby achieving multi-scale semantic fusion. Additionally, to effectively integrate the significant information, we develop a cross-modal fusion strategy (CMS) based on a covariance mask to maximise the complementary properties between different modalities and to suppress the features with high correlation, which are deemed redundant. The experimental results demonstrate that our network outperforms SOTA methods both qualitatively and quantitatively on multiple image fusion benchmarks. The codes are available at https://github.com/Shaoyun2023.
title GrFormer: A Novel Transformer on Grassmann Manifold for Infrared and Visible Image Fusion
topic Computer Vision and Pattern Recognition
I.4
url https://arxiv.org/abs/2506.14384