VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917815046373376 |
|---|---|
| author | Huang, Siyuan Prabhakar, Ram Guo, Yuxiang Chellappa, Rama Peng, Cheng |
| author_facet | Huang, Siyuan Prabhakar, Ram Guo, Yuxiang Chellappa, Rama Peng, Cheng |
| contents | Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and imperfect feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_17074 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification Huang, Siyuan Prabhakar, Ram Guo, Yuxiang Chellappa, Rama Peng, Cheng Computer Vision and Pattern Recognition Person Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and imperfect feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods. |
| title | VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2311.17074 |