Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Qi, Chen, Jixuan, Zhang, Kaiyi, Yu, Xinquan, Chan, Antoni B., Huang, Hui
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913051043692544
author Zhang, Qi
Chen, Jixuan
Zhang, Kaiyi
Yu, Xinquan
Chan, Antoni B.
Huang, Hui
author_facet Zhang, Qi
Chen, Jixuan
Zhang, Kaiyi
Yu, Xinquan
Chan, Antoni B.
Huang, Hui
contents Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking architectures, and most of them are evaluated and compared on relatively small datasets, such as Wildtrack and MultiviewX. Since these two datasets are collected in small scenes and only contain tens of frames in the evaluation stage, it is difficult for the current methods to be applied to real-world applications where scene size and occlusion are more complicated. In this paper, we propose a Transformer-based multi-view crowd tracking model, \textit{MVTrackTrans}, which adopts interactions between camera views and the ground plane for enhanced multi-view tracking performance. Besides, for better evaluation, we collect and label two large real-world multi-view tracking datasets, MVCrowdTrack and CityTrack, which contain a much larger scene size over a longer time period. Compared with existing methods on the two large and new datasets, the proposed MVTrackTrans model achieves better performance, demonstrating the advantages of the model design in dealing with large scenes. We believe the proposed datasets and model will push the frontiers of the task to more practical scenarios, and the datasets and code are available at: https://github.com/zqyq/MVTrackTrans.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19318
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
Zhang, Qi
Chen, Jixuan
Zhang, Kaiyi
Yu, Xinquan
Chan, Antoni B.
Huang, Hui
Computer Vision and Pattern Recognition
Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking architectures, and most of them are evaluated and compared on relatively small datasets, such as Wildtrack and MultiviewX. Since these two datasets are collected in small scenes and only contain tens of frames in the evaluation stage, it is difficult for the current methods to be applied to real-world applications where scene size and occlusion are more complicated. In this paper, we propose a Transformer-based multi-view crowd tracking model, \textit{MVTrackTrans}, which adopts interactions between camera views and the ground plane for enhanced multi-view tracking performance. Besides, for better evaluation, we collect and label two large real-world multi-view tracking datasets, MVCrowdTrack and CityTrack, which contain a much larger scene size over a longer time period. Compared with existing methods on the two large and new datasets, the proposed MVTrackTrans model achieves better performance, demonstrating the advantages of the model design in dealing with large scenes. We believe the proposed datasets and model will push the frontiers of the task to more practical scenarios, and the datasets and code are available at: https://github.com/zqyq/MVTrackTrans.
title Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.19318