Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khanchi, Milad, Amer, Maria, Poullis, Charalambos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913869747716096
author Khanchi, Milad
Amer, Maria
Poullis, Charalambos
author_facet Khanchi, Milad
Amer, Maria
Poullis, Charalambos
contents Current motion-based multiple object tracking (MOT) approaches rely heavily on Intersection-over-Union (IoU) for object association. Without using 3D features, they are ineffective in scenarios with occlusions or visually similar objects. To address this, our paper presents a novel depth-aware framework for MOT. We estimate depth using a zero-shot approach and incorporate it as an independent feature in the association process. Additionally, we introduce a Hierarchical Alignment Score that refines IoU by integrating both coarse bounding box overlap and fine-grained (pixel-level) alignment to improve association accuracy without requiring additional learnable parameters. To our knowledge, this is the first MOT framework to incorporate 3D features (monocular depth) as an independent decision matrix in the association step. Our framework achieves state-of-the-art results on challenging benchmarks without any training nor fine-tuning. The code is available at https://github.com/Milad-Khanchi/DepthMOT
format Preprint
id arxiv_https___arxiv_org_abs_2506_00774
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking
Khanchi, Milad
Amer, Maria
Poullis, Charalambos
Computer Vision and Pattern Recognition
Current motion-based multiple object tracking (MOT) approaches rely heavily on Intersection-over-Union (IoU) for object association. Without using 3D features, they are ineffective in scenarios with occlusions or visually similar objects. To address this, our paper presents a novel depth-aware framework for MOT. We estimate depth using a zero-shot approach and incorporate it as an independent feature in the association process. Additionally, we introduce a Hierarchical Alignment Score that refines IoU by integrating both coarse bounding box overlap and fine-grained (pixel-level) alignment to improve association accuracy without requiring additional learnable parameters. To our knowledge, this is the first MOT framework to incorporate 3D features (monocular depth) as an independent decision matrix in the association step. Our framework achieves state-of-the-art results on challenging benchmarks without any training nor fine-tuning. The code is available at https://github.com/Milad-Khanchi/DepthMOT
title Depth-Aware Scoring and Hierarchical Alignment for Multiple Object Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.00774