Text-based Aerial-Ground Person Retrieval

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Xinyu, Wu, Yu, Ma, Jiayao, Wang, Wenhao, Cao, Min, Ye, Mang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912702211817472
author Zhou, Xinyu
Wu, Yu
Ma, Jiayao
Wang, Wenhao
Cao, Min
Ye, Mang
author_facet Zhou, Xinyu
Wu, Yu
Ma, Jiayao
Wang, Wenhao
Cao, Min
Ye, Mang
contents This work introduces Text-based Aerial-Ground Person Retrieval (TAG-PR), which aims to retrieve person images from heterogeneous aerial and ground views with textual descriptions. Unlike traditional Text-based Person Retrieval (T-PR), which focuses solely on ground-view images, TAG-PR introduces greater practical significance and presents unique challenges due to the large viewpoint discrepancy across images. To support this task, we contribute: (1) TAG-PEDES dataset, constructed from public benchmarks with automatically generated textual descriptions, enhanced by a diversified text generation paradigm to ensure robustness under view heterogeneity; and (2) TAG-CLIP, a novel retrieval framework that addresses view heterogeneity through a hierarchically-routed mixture of experts module to learn view-specific and view-agnostic features and a viewpoint decoupling strategy to decouple view-specific features for better cross-modal alignment. We evaluate the effectiveness of TAG-CLIP on both the proposed TAG-PEDES dataset and existing T-PR benchmarks. The dataset and code are available at https://github.com/Flame-Chasers/TAG-PR.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08369
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-based Aerial-Ground Person Retrieval
Zhou, Xinyu
Wu, Yu
Ma, Jiayao
Wang, Wenhao
Cao, Min
Ye, Mang
Computer Vision and Pattern Recognition
Artificial Intelligence
This work introduces Text-based Aerial-Ground Person Retrieval (TAG-PR), which aims to retrieve person images from heterogeneous aerial and ground views with textual descriptions. Unlike traditional Text-based Person Retrieval (T-PR), which focuses solely on ground-view images, TAG-PR introduces greater practical significance and presents unique challenges due to the large viewpoint discrepancy across images. To support this task, we contribute: (1) TAG-PEDES dataset, constructed from public benchmarks with automatically generated textual descriptions, enhanced by a diversified text generation paradigm to ensure robustness under view heterogeneity; and (2) TAG-CLIP, a novel retrieval framework that addresses view heterogeneity through a hierarchically-routed mixture of experts module to learn view-specific and view-agnostic features and a viewpoint decoupling strategy to decouple view-specific features for better cross-modal alignment. We evaluate the effectiveness of TAG-CLIP on both the proposed TAG-PEDES dataset and existing T-PR benchmarks. The dataset and code are available at https://github.com/Flame-Chasers/TAG-PR.
title Text-based Aerial-Ground Person Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.08369