Frequency Domain Modality-invariant Feature Learning for Visible-infrared Person Re-Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yulin, Zhang, Tianzhu, Zhang, Yongdong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914629412716544
author Li, Yulin
Zhang, Tianzhu
Zhang, Yongdong
author_facet Li, Yulin
Zhang, Tianzhu
Zhang, Yongdong
contents Visible-infrared person re-identification (VI-ReID) is challenging due to the significant cross-modality discrepancies between visible and infrared images. While existing methods have focused on designing complex network architectures or using metric learning constraints to learn modality-invariant features, they often overlook which specific component of the image causes the modality discrepancy problem. In this paper, we first reveal that the difference in the amplitude component of visible and infrared images is the primary factor that causes the modality discrepancy and further propose a novel Frequency Domain modality-invariant feature learning framework (FDMNet) to reduce modality discrepancy from the frequency domain perspective. Our framework introduces two novel modules, namely the Instance-Adaptive Amplitude Filter (IAF) module and the Phrase-Preserving Normalization (PPNorm) module, to enhance the modality-invariant amplitude component and suppress the modality-specific component at both the image- and feature-levels. Extensive experimental results on two standard benchmarks, SYSU-MM01 and RegDB, demonstrate the superior performance of our FDMNet against state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2401_01839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Frequency Domain Modality-invariant Feature Learning for Visible-infrared Person Re-Identification
Li, Yulin
Zhang, Tianzhu
Zhang, Yongdong
Computer Vision and Pattern Recognition
Visible-infrared person re-identification (VI-ReID) is challenging due to the significant cross-modality discrepancies between visible and infrared images. While existing methods have focused on designing complex network architectures or using metric learning constraints to learn modality-invariant features, they often overlook which specific component of the image causes the modality discrepancy problem. In this paper, we first reveal that the difference in the amplitude component of visible and infrared images is the primary factor that causes the modality discrepancy and further propose a novel Frequency Domain modality-invariant feature learning framework (FDMNet) to reduce modality discrepancy from the frequency domain perspective. Our framework introduces two novel modules, namely the Instance-Adaptive Amplitude Filter (IAF) module and the Phrase-Preserving Normalization (PPNorm) module, to enhance the modality-invariant amplitude component and suppress the modality-specific component at both the image- and feature-levels. Extensive experimental results on two standard benchmarks, SYSU-MM01 and RegDB, demonstrate the superior performance of our FDMNet against state-of-the-art methods.
title Frequency Domain Modality-invariant Feature Learning for Visible-infrared Person Re-Identification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.01839