Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Lanyun, Chen, Tianrun, Ji, Deyi, Ye, Jieping, Liu, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916586564091904
author Zhu, Lanyun
Chen, Tianrun
Ji, Deyi
Ye, Jieping
Liu, Jun
author_facet Zhu, Lanyun
Chen, Tianrun
Ji, Deyi
Ye, Jieping
Liu, Jun
contents This paper proposes a new effective and efficient plug-and-play backbone for video-based person re-identification (ReID). Conventional video-based ReID methods typically use CNN or transformer backbones to extract deep features for every position in every sampled video frame. Here, we argue that this exhaustive feature extraction could be unnecessary, since we find that different frames in a ReID video often exhibit small differences and contain many similar regions due to the relatively slight movements of human beings. Inspired by this, a more selective, efficient paradigm is explored in this paper. Specifically, we introduce a patch selection mechanism to reduce computational cost by choosing only the crucial and non-repetitive patches for feature extraction. Additionally, we present a novel network structure that generates and utilizes pseudo frame global context to address the issue of incomplete views resulting from sparse inputs. By incorporating these new designs, our backbone can achieve both high performance and low computational cost. Extensive experiments on multiple datasets show that our approach reduces the computational cost by 74\% compared to ViT-B and 28\% compared to ResNet50, while the accuracy is on par with ViT-B and outperforms ResNet50 significantly.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
Zhu, Lanyun
Chen, Tianrun
Ji, Deyi
Ye, Jieping
Liu, Jun
Computer Vision and Pattern Recognition
This paper proposes a new effective and efficient plug-and-play backbone for video-based person re-identification (ReID). Conventional video-based ReID methods typically use CNN or transformer backbones to extract deep features for every position in every sampled video frame. Here, we argue that this exhaustive feature extraction could be unnecessary, since we find that different frames in a ReID video often exhibit small differences and contain many similar regions due to the relatively slight movements of human beings. Inspired by this, a more selective, efficient paradigm is explored in this paper. Specifically, we introduce a patch selection mechanism to reduce computational cost by choosing only the crucial and non-repetitive patches for feature extraction. Additionally, we present a novel network structure that generates and utilizes pseudo frame global context to address the issue of incomplete views resulting from sparse inputs. By incorporating these new designs, our backbone can achieve both high performance and low computational cost. Extensive experiments on multiple datasets show that our approach reduces the computational cost by 74\% compared to ViT-B and 28\% compared to ResNet50, while the accuracy is on par with ViT-B and outperforms ResNet50 significantly.
title Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.16811