VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Siyuan, Prabhakar, Ram, Guo, Yuxiang, Chellappa, Rama, Peng, Cheng
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917815046373376
author Huang, Siyuan
Prabhakar, Ram
Guo, Yuxiang
Chellappa, Rama
Peng, Cheng
author_facet Huang, Siyuan
Prabhakar, Ram
Guo, Yuxiang
Chellappa, Rama
Peng, Cheng
contents Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and imperfect feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17074
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification
Huang, Siyuan
Prabhakar, Ram
Guo, Yuxiang
Chellappa, Rama
Peng, Cheng
Computer Vision and Pattern Recognition
Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and imperfect feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods.
title VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.17074