JDT3D: Addressing the Gaps in LiDAR-Based Tracking-by-Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheong, Brian, Zhou, Jiachen, Waslander, Steven
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917722671022080
author Cheong, Brian
Zhou, Jiachen
Waslander, Steven
author_facet Cheong, Brian
Zhou, Jiachen
Waslander, Steven
contents Tracking-by-detection (TBD) methods achieve state-of-the-art performance on 3D tracking benchmarks for autonomous driving. On the other hand, tracking-by-attention (TBA) methods have the potential to outperform TBD methods, particularly for long occlusions and challenging detection settings. This work investigates why TBA methods continue to lag in performance behind TBD methods using a LiDAR-based joint detector and tracker called JDT3D. Based on this analysis, we propose two generalizable methods to bridge the gap between TBD and TBA methods: track sampling augmentation and confidence-based query propagation. JDT3D is trained and evaluated on the nuScenes dataset, achieving 0.574 on the AMOTA metric on the nuScenes test set, outperforming all existing LiDAR-based TBA approaches by over 6%. Based on our results, we further discuss some potential challenges with the existing TBA model formulation to explain the continued gap in performance with TBD methods. The implementation of JDT3D can be found at the following link: https://github.com/TRAILab/JDT3D.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04926
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle JDT3D: Addressing the Gaps in LiDAR-Based Tracking-by-Attention
Cheong, Brian
Zhou, Jiachen
Waslander, Steven
Computer Vision and Pattern Recognition
Tracking-by-detection (TBD) methods achieve state-of-the-art performance on 3D tracking benchmarks for autonomous driving. On the other hand, tracking-by-attention (TBA) methods have the potential to outperform TBD methods, particularly for long occlusions and challenging detection settings. This work investigates why TBA methods continue to lag in performance behind TBD methods using a LiDAR-based joint detector and tracker called JDT3D. Based on this analysis, we propose two generalizable methods to bridge the gap between TBD and TBA methods: track sampling augmentation and confidence-based query propagation. JDT3D is trained and evaluated on the nuScenes dataset, achieving 0.574 on the AMOTA metric on the nuScenes test set, outperforming all existing LiDAR-based TBA approaches by over 6%. Based on our results, we further discuss some potential challenges with the existing TBA model formulation to explain the continued gap in performance with TBD methods. The implementation of JDT3D can be found at the following link: https://github.com/TRAILab/JDT3D.
title JDT3D: Addressing the Gaps in LiDAR-Based Tracking-by-Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.04926