3D Ground Truth Reconstruction from Multi-Camera Annotations Using UKF

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Van Ma, Linh, Fatima, Unse, Chriv, Tepy Sokun, Imran, Haroon, Jeon, Moongu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918214714261504
author Van Ma, Linh
Fatima, Unse
Chriv, Tepy Sokun
Imran, Haroon
Jeon, Moongu
author_facet Van Ma, Linh
Fatima, Unse
Chriv, Tepy Sokun
Imran, Haroon
Jeon, Moongu
contents Accurate 3D ground truth estimation is critical for applications such as autonomous navigation, surveillance, and robotics. This paper introduces a novel method that uses an Unscented Kalman Filter (UKF) to fuse 2D bounding box or pose keypoint ground truth annotations from multiple calibrated cameras into accurate 3D ground truth. By leveraging human-annotated ground-truth 2D, our proposed method, a multi-camera single-object tracking algorithm, transforms 2D image coordinates into robust 3D world coordinates through homography-based projection and UKF-based fusion. Our proposed algorithm processes multi-view data to estimate object positions and shapes while effectively handling challenges such as occlusion. We evaluate our method on the CMC, Wildtrack, and Panoptic datasets, demonstrating high accuracy in 3D localization compared to the available 3D ground truth. Unlike existing approaches that provide only ground-plane information, our method also outputs the full 3D shape of each object. Additionally, the algorithm offers a scalable and fully automatic solution for multi-camera systems using only 2D image annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17609
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D Ground Truth Reconstruction from Multi-Camera Annotations Using UKF
Van Ma, Linh
Fatima, Unse
Chriv, Tepy Sokun
Imran, Haroon
Jeon, Moongu
Computer Vision and Pattern Recognition
Accurate 3D ground truth estimation is critical for applications such as autonomous navigation, surveillance, and robotics. This paper introduces a novel method that uses an Unscented Kalman Filter (UKF) to fuse 2D bounding box or pose keypoint ground truth annotations from multiple calibrated cameras into accurate 3D ground truth. By leveraging human-annotated ground-truth 2D, our proposed method, a multi-camera single-object tracking algorithm, transforms 2D image coordinates into robust 3D world coordinates through homography-based projection and UKF-based fusion. Our proposed algorithm processes multi-view data to estimate object positions and shapes while effectively handling challenges such as occlusion. We evaluate our method on the CMC, Wildtrack, and Panoptic datasets, demonstrating high accuracy in 3D localization compared to the available 3D ground truth. Unlike existing approaches that provide only ground-plane information, our method also outputs the full 3D shape of each object. Additionally, the algorithm offers a scalable and fully automatic solution for multi-camera systems using only 2D image annotations.
title 3D Ground Truth Reconstruction from Multi-Camera Annotations Using UKF
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17609