AAformer: Auto-Aligned Transformer for Person Re-Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Kuan, Guo, Haiyun, Zhang, Shiliang, Wang, Yaowei, Liu, Jing, Wang, Jinqiao, Tang, Ming
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913402907000832
author Zhu, Kuan
Guo, Haiyun
Zhang, Shiliang
Wang, Yaowei
Liu, Jing
Wang, Jinqiao
Tang, Ming
author_facet Zhu, Kuan
Guo, Haiyun
Zhang, Shiliang
Wang, Yaowei
Liu, Jing
Wang, Jinqiao
Tang, Ming
contents In person re-identification (re-ID), extracting part-level features from person images has been verified to be crucial to offer fine-grained information. Most of the existing CNN-based methods only locate the human parts coarsely, or rely on pretrained human parsing models and fail in locating the identifiable nonhuman parts (e.g., knapsack). In this article, we introduce an alignment scheme in transformer architecture for the first time and propose the auto-aligned transformer (AAformer) to automatically locate both the human parts and nonhuman ones at patch level. We introduce the "Part tokens ([PART]s)", which are learnable vectors, to extract part features in the transformer. A [PART] only interacts with a local subset of patches in self-attention and learns to be the part representation. To adaptively group the image patches into different subsets, we design the auto-alignment. Auto-alignment employs a fast variant of optimal transport (OT) algorithm to online cluster the patch embeddings into several groups with the [PART]s as their prototypes. AAformer integrates the part alignment into the self-attention and the output [PART]s can be directly used as part features for retrieval. Extensive experiments validate the effectiveness of [PART]s and the superiority of AAformer over various state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2104_00921
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle AAformer: Auto-Aligned Transformer for Person Re-Identification
Zhu, Kuan
Guo, Haiyun
Zhang, Shiliang
Wang, Yaowei
Liu, Jing
Wang, Jinqiao
Tang, Ming
Computer Vision and Pattern Recognition
In person re-identification (re-ID), extracting part-level features from person images has been verified to be crucial to offer fine-grained information. Most of the existing CNN-based methods only locate the human parts coarsely, or rely on pretrained human parsing models and fail in locating the identifiable nonhuman parts (e.g., knapsack). In this article, we introduce an alignment scheme in transformer architecture for the first time and propose the auto-aligned transformer (AAformer) to automatically locate both the human parts and nonhuman ones at patch level. We introduce the "Part tokens ([PART]s)", which are learnable vectors, to extract part features in the transformer. A [PART] only interacts with a local subset of patches in self-attention and learns to be the part representation. To adaptively group the image patches into different subsets, we design the auto-alignment. Auto-alignment employs a fast variant of optimal transport (OT) algorithm to online cluster the patch embeddings into several groups with the [PART]s as their prototypes. AAformer integrates the part alignment into the self-attention and the output [PART]s can be directly used as part features for retrieval. Extensive experiments validate the effectiveness of [PART]s and the superiority of AAformer over various state-of-the-art methods.
title AAformer: Auto-Aligned Transformer for Person Re-Identification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2104.00921