Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Qingyu, Wang, Longguang, Sheng, Weidong, Wang, Yingqian, Xiao, Chao, Ma, Chao, An, Wei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915063783227392
author Xu, Qingyu
Wang, Longguang
Sheng, Weidong
Wang, Yingqian
Xiao, Chao
Ma, Chao
An, Wei
author_facet Xu, Qingyu
Wang, Longguang
Sheng, Weidong
Wang, Yingqian
Xiao, Chao
Ma, Chao
An, Wei
contents Tracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on single-modality scenes, and overlook the complementary characteristics of tiny objects captured by multiple remote sensors. To enhance tracking performance by integrating complementary information from multiple sources, we propose a novel framework called {HGT-Track (Heterogeneous Graph Transformer based Multi-Tiny-Object Tracking)}. Specifically, we first employ a Transformer-based encoder to embed images from different modalities. Subsequently, we utilize Heterogeneous Graph Transformer to aggregate spatial and temporal information from multiple modalities to generate detection and tracking features. Additionally, we introduce a target re-detection module (ReDet) to ensure tracklet continuity by maintaining consistency across different modalities. Furthermore, this paper introduces the first benchmark VT-Tiny-MOT (Visible-Thermal Tiny Multi-Object Tracking) for RGB-T fused multiple tiny object tracking. Extensive experiments are conducted on VT-Tiny-MOT, and the results have demonstrated the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of MOTA (Multiple-Object Tracking Accuracy) and ID-F1 score. The code and dataset will be made available at https://github.com/xuqingyu26/HGTMT.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10861
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos
Xu, Qingyu
Wang, Longguang
Sheng, Weidong
Wang, Yingqian
Xiao, Chao
Ma, Chao
An, Wei
Computer Vision and Pattern Recognition
Artificial Intelligence
Tracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on single-modality scenes, and overlook the complementary characteristics of tiny objects captured by multiple remote sensors. To enhance tracking performance by integrating complementary information from multiple sources, we propose a novel framework called {HGT-Track (Heterogeneous Graph Transformer based Multi-Tiny-Object Tracking)}. Specifically, we first employ a Transformer-based encoder to embed images from different modalities. Subsequently, we utilize Heterogeneous Graph Transformer to aggregate spatial and temporal information from multiple modalities to generate detection and tracking features. Additionally, we introduce a target re-detection module (ReDet) to ensure tracklet continuity by maintaining consistency across different modalities. Furthermore, this paper introduces the first benchmark VT-Tiny-MOT (Visible-Thermal Tiny Multi-Object Tracking) for RGB-T fused multiple tiny object tracking. Extensive experiments are conducted on VT-Tiny-MOT, and the results have demonstrated the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of MOTA (Multiple-Object Tracking Accuracy) and ID-F1 score. The code and dataset will be made available at https://github.com/xuqingyu26/HGTMT.
title Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.10861