MCBLT: Multi-Camera Multi-Object 3D Tracking in Long Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yizhou, Meinhardt, Tim, Cetintas, Orcun, Yang, Cheng-Yen, Pusegaonkar, Sameer Satish, Missaoui, Benjamin, Biswas, Sujit, Tang, Zheng, Leal-Taixé, Laura
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910894966964224
author Wang, Yizhou
Meinhardt, Tim
Cetintas, Orcun
Yang, Cheng-Yen
Pusegaonkar, Sameer Satish
Missaoui, Benjamin
Biswas, Sujit
Tang, Zheng
Leal-Taixé, Laura
author_facet Wang, Yizhou
Meinhardt, Tim
Cetintas, Orcun
Yang, Cheng-Yen
Pusegaonkar, Sameer Satish
Missaoui, Benjamin
Biswas, Sujit
Tang, Zheng
Leal-Taixé, Laura
contents Object perception from multi-view cameras is crucial for intelligent systems, particularly in indoor environments, e.g., warehouses, retail stores, and hospitals. Most traditional multi-target multi-camera (MTMC) detection and tracking methods rely on 2D object detection, single-view multi-object tracking (MOT), and cross-view re-identification (ReID) techniques, without properly handling important 3D information by multi-view image aggregation. In this paper, we propose a 3D object detection and tracking framework, named MCBLT, which first aggregates multi-view images with necessary camera calibration parameters to obtain 3D object detections in bird's-eye view (BEV). Then, we introduce hierarchical graph neural networks (GNNs) to track these 3D detections in BEV for MTMC tracking results. Unlike existing methods, MCBLT has impressive generalizability across different scenes and diverse camera settings, with exceptional capability for long-term association handling. As a result, our proposed MCBLT establishes a new state-of-the-art on the AICity'24 dataset with $81.22$ HOTA, and on the WildTrack dataset with $95.6$ IDF1.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00692
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MCBLT: Multi-Camera Multi-Object 3D Tracking in Long Videos
Wang, Yizhou
Meinhardt, Tim
Cetintas, Orcun
Yang, Cheng-Yen
Pusegaonkar, Sameer Satish
Missaoui, Benjamin
Biswas, Sujit
Tang, Zheng
Leal-Taixé, Laura
Computer Vision and Pattern Recognition
Object perception from multi-view cameras is crucial for intelligent systems, particularly in indoor environments, e.g., warehouses, retail stores, and hospitals. Most traditional multi-target multi-camera (MTMC) detection and tracking methods rely on 2D object detection, single-view multi-object tracking (MOT), and cross-view re-identification (ReID) techniques, without properly handling important 3D information by multi-view image aggregation. In this paper, we propose a 3D object detection and tracking framework, named MCBLT, which first aggregates multi-view images with necessary camera calibration parameters to obtain 3D object detections in bird's-eye view (BEV). Then, we introduce hierarchical graph neural networks (GNNs) to track these 3D detections in BEV for MTMC tracking results. Unlike existing methods, MCBLT has impressive generalizability across different scenes and diverse camera settings, with exceptional capability for long-term association handling. As a result, our proposed MCBLT establishes a new state-of-the-art on the AICity'24 dataset with $81.22$ HOTA, and on the WildTrack dataset with $95.6$ IDF1.
title MCBLT: Multi-Camera Multi-Object 3D Tracking in Long Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00692