Exploring Part-Informed Visual-Language Learning for Person Re-Identification
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912284359524352 |
|---|---|
| author | Lin, Yin Chen, Yehansen Yin, Baocai Hu, Jinshui Yin, Bing Liu, Cong Wang, Zengfu |
| author_facet | Lin, Yin Chen, Yehansen Yin, Baocai Hu, Jinshui Yin, Bing Liu, Cong Wang, Zengfu |
| contents | Recently, visual-language learning (VLL) has shown great potential in enhancing visual-based person re-identification (ReID). Existing VLL-based ReID methods typically focus on image-text feature alignment at the whole-body level, while neglecting supervision on fine-grained part features, thus lacking constraints for local feature semantic consistency. To this end, we propose Part-Informed Visual-language Learning ($π$-VL) to enhance fine-grained visual features with part-informed language supervisions for ReID tasks. Specifically, $π$-VL introduces a human parsing-guided prompt tuning strategy and a hierarchical visual-language alignment paradigm to ensure within-part feature semantic consistency. The former combines both identity labels and human parsing maps to constitute pixel-level text prompts, and the latter fuses multi-scale visual features with a light-weight auxiliary head to perform fine-grained image-text alignment. As a plug-and-play and inference-free solution, our $π$-VL achieves performance comparable to or better than state-of-the-art methods on four commonly used ReID benchmarks. Notably, it reports 91.0% Rank-1 and 76.9% mAP on the challenging MSMT17 database, without bells and whistles. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2308_02738 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Exploring Part-Informed Visual-Language Learning for Person Re-Identification Lin, Yin Chen, Yehansen Yin, Baocai Hu, Jinshui Yin, Bing Liu, Cong Wang, Zengfu Computer Vision and Pattern Recognition Recently, visual-language learning (VLL) has shown great potential in enhancing visual-based person re-identification (ReID). Existing VLL-based ReID methods typically focus on image-text feature alignment at the whole-body level, while neglecting supervision on fine-grained part features, thus lacking constraints for local feature semantic consistency. To this end, we propose Part-Informed Visual-language Learning ($π$-VL) to enhance fine-grained visual features with part-informed language supervisions for ReID tasks. Specifically, $π$-VL introduces a human parsing-guided prompt tuning strategy and a hierarchical visual-language alignment paradigm to ensure within-part feature semantic consistency. The former combines both identity labels and human parsing maps to constitute pixel-level text prompts, and the latter fuses multi-scale visual features with a light-weight auxiliary head to perform fine-grained image-text alignment. As a plug-and-play and inference-free solution, our $π$-VL achieves performance comparable to or better than state-of-the-art methods on four commonly used ReID benchmarks. Notably, it reports 91.0% Rank-1 and 76.9% mAP on the challenging MSMT17 database, without bells and whistles. |
| title | Exploring Part-Informed Visual-Language Learning for Person Re-Identification |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2308.02738 |