PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Lei, Dai, Pingyang, Chen, Jie, Cao, Liujuan, Wu, Yongjian, Ji, Rongrong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909300576747520
author Tan, Lei
Dai, Pingyang
Chen, Jie
Cao, Liujuan
Wu, Yongjian
Ji, Rongrong
author_facet Tan, Lei
Dai, Pingyang
Chen, Jie
Cao, Liujuan
Wu, Yongjian
Ji, Rongrong
contents Extracting robust feature representation is critical for object re-identification to accurately identify objects across non-overlapping cameras. Although having a strong representation ability, the Vision Transformer (ViT) tends to overfit on most distinct regions of training data, limiting its generalizability and attention to holistic object features. Meanwhile, due to the structural difference between CNN and ViT, fine-grained strategies that effectively address this issue in CNN do not continue to be successful in ViT. To address this issue, by observing the latent diverse representation hidden behind the multi-head attention, we present PartFormer, an innovative adaptation of ViT designed to overcome the granularity limitations in object Re-ID tasks. The PartFormer integrates a Head Disentangling Block (HDB) that awakens the diverse representation of multi-head self-attention without the typical loss of feature richness induced by concatenation and FFN layers post-attention. To avoid the homogenization of attention heads and promote robust part-based feature learning, two head diversity constraints are imposed: attention diversity constraint and correlation diversity constraint. These constraints enable the model to exploit diverse and discriminative feature representations from different attention heads. Comprehensive experiments on various object Re-ID benchmarks demonstrate the superiority of the PartFormer. Specifically, our framework significantly outperforms state-of-the-art by 2.4\% mAP scores on the most challenging MSMT17 dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16684
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification
Tan, Lei
Dai, Pingyang
Chen, Jie
Cao, Liujuan
Wu, Yongjian
Ji, Rongrong
Computer Vision and Pattern Recognition
Extracting robust feature representation is critical for object re-identification to accurately identify objects across non-overlapping cameras. Although having a strong representation ability, the Vision Transformer (ViT) tends to overfit on most distinct regions of training data, limiting its generalizability and attention to holistic object features. Meanwhile, due to the structural difference between CNN and ViT, fine-grained strategies that effectively address this issue in CNN do not continue to be successful in ViT. To address this issue, by observing the latent diverse representation hidden behind the multi-head attention, we present PartFormer, an innovative adaptation of ViT designed to overcome the granularity limitations in object Re-ID tasks. The PartFormer integrates a Head Disentangling Block (HDB) that awakens the diverse representation of multi-head self-attention without the typical loss of feature richness induced by concatenation and FFN layers post-attention. To avoid the homogenization of attention heads and promote robust part-based feature learning, two head diversity constraints are imposed: attention diversity constraint and correlation diversity constraint. These constraints enable the model to exploit diverse and discriminative feature representations from different attention heads. Comprehensive experiments on various object Re-ID benchmarks demonstrate the superiority of the PartFormer. Specifically, our framework significantly outperforms state-of-the-art by 2.4\% mAP scores on the most challenging MSMT17 dataset.
title PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.16684