Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Jiayi, Tang, Shengeng, Liu, Ao, Cheng, Lechao, Wu, Jingjing, Wei, Yanyan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915139708518400
author He, Jiayi
Tang, Shengeng
Liu, Ao
Cheng, Lechao
Wu, Jingjing
Wei, Yanyan
author_facet He, Jiayi
Tang, Shengeng
Liu, Ao
Cheng, Lechao
Wu, Jingjing
Wei, Yanyan
contents This paper presents the HFUT-LMC team's solution to the WWW 2025 challenge on Text-based Person Anomaly Search (TPAS). The primary objective of this challenge is to accurately identify pedestrians exhibiting either normal or abnormal behavior within a large library of pedestrian images. Unlike traditional video analysis tasks, TPAS significantly emphasizes understanding and interpreting the subtle relationships between text descriptions and visual data. The complexity of this task lies in the model's need to not only match individuals to text descriptions in massive image datasets but also accurately differentiate between search results when faced with similar descriptions. To overcome these challenges, we introduce the Similarity Coverage Analysis (SCA) strategy to address the recognition difficulty caused by similar text descriptions. This strategy effectively enhances the model's capacity to manage subtle differences, thus improving both the accuracy and reliability of the search. Our proposed solution demonstrated excellent performance in this challenge.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
He, Jiayi
Tang, Shengeng
Liu, Ao
Cheng, Lechao
Wu, Jingjing
Wei, Yanyan
Computer Vision and Pattern Recognition
Multimedia
This paper presents the HFUT-LMC team's solution to the WWW 2025 challenge on Text-based Person Anomaly Search (TPAS). The primary objective of this challenge is to accurately identify pedestrians exhibiting either normal or abnormal behavior within a large library of pedestrian images. Unlike traditional video analysis tasks, TPAS significantly emphasizes understanding and interpreting the subtle relationships between text descriptions and visual data. The complexity of this task lies in the model's need to not only match individuals to text descriptions in massive image datasets but also accurately differentiate between search results when faced with similar descriptions. To overcome these challenges, we introduce the Similarity Coverage Analysis (SCA) strategy to address the recognition difficulty caused by similar text descriptions. This strategy effectively enhances the model's capacity to manage subtle differences, thus improving both the accuracy and reliability of the search. Our proposed solution demonstrated excellent performance in this challenge.
title Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2502.03230