ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Can, Huang, Siteng, Song, Wenxuan, Ding, Pengxiang, Zhang, Min, Wang, Donglin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917790546395136
author Cui, Can
Huang, Siteng
Song, Wenxuan
Ding, Pengxiang
Zhang, Min
Wang, Donglin
author_facet Cui, Can
Huang, Siteng
Song, Wenxuan
Ding, Pengxiang
Zhang, Min
Wang, Donglin
contents To address the occlusion issues in person Re-Identification (ReID) tasks, many methods have been proposed to extract part features by introducing external spatial information. However, due to missing part appearance information caused by occlusion and noisy spatial information from external model, these purely vision-based approaches fail to correctly learn the features of human body parts from limited training data and struggle in accurately locating body parts, ultimately leading to misaligned part features. To tackle these challenges, we propose a Prompt-guided Feature Disentangling method (ProFD), which leverages the rich pre-trained knowledge in the textual modality facilitate model to generate well-aligned part features. ProFD first designs part-specific prompts and utilizes noisy segmentation mask to preliminarily align visual and textual embedding, enabling the textual prompts to have spatial awareness. Furthermore, to alleviate the noise from external masks, ProFD adopts a hybrid-attention decoder, ensuring spatial and semantic consistency during the decoding process to minimize noise impact. Additionally, to avoid catastrophic forgetting, we employ a self-distillation strategy, retaining pre-trained knowledge of CLIP to mitigate over-fitting. Evaluation results on the Market1501, DukeMTMC-ReID, Occluded-Duke, Occluded-ReID, and P-DukeMTMC datasets demonstrate that ProFD achieves state-of-the-art results. Our project is available at: https://github.com/Cuixxx/ProFD.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
Cui, Can
Huang, Siteng
Song, Wenxuan
Ding, Pengxiang
Zhang, Min
Wang, Donglin
Computer Vision and Pattern Recognition
Multimedia
To address the occlusion issues in person Re-Identification (ReID) tasks, many methods have been proposed to extract part features by introducing external spatial information. However, due to missing part appearance information caused by occlusion and noisy spatial information from external model, these purely vision-based approaches fail to correctly learn the features of human body parts from limited training data and struggle in accurately locating body parts, ultimately leading to misaligned part features. To tackle these challenges, we propose a Prompt-guided Feature Disentangling method (ProFD), which leverages the rich pre-trained knowledge in the textual modality facilitate model to generate well-aligned part features. ProFD first designs part-specific prompts and utilizes noisy segmentation mask to preliminarily align visual and textual embedding, enabling the textual prompts to have spatial awareness. Furthermore, to alleviate the noise from external masks, ProFD adopts a hybrid-attention decoder, ensuring spatial and semantic consistency during the decoding process to minimize noise impact. Additionally, to avoid catastrophic forgetting, we employ a self-distillation strategy, retaining pre-trained knowledge of CLIP to mitigate over-fitting. Evaluation results on the Market1501, DukeMTMC-ReID, Occluded-Duke, Occluded-ReID, and P-DukeMTMC datasets demonstrate that ProFD achieves state-of-the-art results. Our project is available at: https://github.com/Cuixxx/ProFD.
title ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2409.20081