Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Haonan, Wang, Xinyao, Wu, Boxi, Zheng, Tu, Yunhua, Wang, Yang, Zheng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917120594411520
author Zhang, Haonan
Wang, Xinyao
Wu, Boxi
Zheng, Tu
Yunhua, Wang
Yang, Zheng
author_facet Zhang, Haonan
Wang, Xinyao
Wu, Boxi
Zheng, Tu
Yunhua, Wang
Yang, Zheng
contents 3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters, to predict trajectories. While effective in simple scenarios, this approach often struggles in crowded environments or with inaccurate detections, as it overlooks the rich geometric relationships between objects. This highlights the need to leverage spatial cues. However, existing geometry-aware methods can be susceptible to interference from irrelevant objects, leading to ambiguous features and incorrect associations. To address this, we propose focusing on cue-consistency: identifying and matching stable spatial patterns over time. We introduce the Dynamic Scene Cue-Consistency Tracker (DSC-Track) to implement this principle. Firstly, we design a unified spatiotemporal encoder using Point Pair Features (PPF) to learn discriminative trajectory embeddings while suppressing interference. Secondly, our cue-consistency transformer module explicitly aligns consistent feature representations between historical tracks and current detections. Finally, a dynamic update mechanism preserves salient spatiotemporal information for stable online tracking. Extensive experiments on the nuScenes and Waymo Open Datasets validate the effectiveness and robustness of our approach. On the nuScenes benchmark, for instance, our method achieves state-of-the-art performance, reaching 73.2% and 70.3% AMOTA on the validation and test sets, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
Zhang, Haonan
Wang, Xinyao
Wu, Boxi
Zheng, Tu
Yunhua, Wang
Yang, Zheng
Computer Vision and Pattern Recognition
3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters, to predict trajectories. While effective in simple scenarios, this approach often struggles in crowded environments or with inaccurate detections, as it overlooks the rich geometric relationships between objects. This highlights the need to leverage spatial cues. However, existing geometry-aware methods can be susceptible to interference from irrelevant objects, leading to ambiguous features and incorrect associations. To address this, we propose focusing on cue-consistency: identifying and matching stable spatial patterns over time. We introduce the Dynamic Scene Cue-Consistency Tracker (DSC-Track) to implement this principle. Firstly, we design a unified spatiotemporal encoder using Point Pair Features (PPF) to learn discriminative trajectory embeddings while suppressing interference. Secondly, our cue-consistency transformer module explicitly aligns consistent feature representations between historical tracks and current detections. Finally, a dynamic update mechanism preserves salient spatiotemporal information for stable online tracking. Extensive experiments on the nuScenes and Waymo Open Datasets validate the effectiveness and robustness of our approach. On the nuScenes benchmark, for instance, our method achieves state-of-the-art performance, reaching 73.2% and 70.3% AMOTA on the validation and test sets, respectively.
title Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.11323