Saved in:
Bibliographic Details
Main Authors: Tushe, Ergi, Farooq, Bilal
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.03541
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912521561047040
author Tushe, Ergi
Farooq, Bilal
author_facet Tushe, Ergi
Farooq, Bilal
contents The integration of Automated Delivery Robots (ADRs) into pedestrian-heavy urban spaces introduces unique challenges in terms of safe, efficient, and socially acceptable navigation. We develop the complete pipeline for a single vision sensor based multi-pedestrian detection and tracking, pose estimation, and monocular depth perception. Leveraging the real-world MOT17 dataset sequences, this study demonstrates how integrating human-pose estimation and depth cues enhances pedestrian trajectory prediction and identity maintenance, even under occlusions and dense crowds. Results show measurable improvements, including up to a 10% increase in identity preservation (IDF1), a 7% improvement in multiobject tracking accuracy (MOTA), and consistently high detection precision exceeding 85%, even in challenging scenarios. Notably, the system identifies vulnerable pedestrian groups supporting more socially aware and inclusive robot behaviour.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-based Perception System for Automated Delivery Robot-Pedestrians Interactions
Tushe, Ergi
Farooq, Bilal
Robotics
Machine Learning
The integration of Automated Delivery Robots (ADRs) into pedestrian-heavy urban spaces introduces unique challenges in terms of safe, efficient, and socially acceptable navigation. We develop the complete pipeline for a single vision sensor based multi-pedestrian detection and tracking, pose estimation, and monocular depth perception. Leveraging the real-world MOT17 dataset sequences, this study demonstrates how integrating human-pose estimation and depth cues enhances pedestrian trajectory prediction and identity maintenance, even under occlusions and dense crowds. Results show measurable improvements, including up to a 10% increase in identity preservation (IDF1), a 7% improvement in multiobject tracking accuracy (MOTA), and consistently high detection precision exceeding 85%, even in challenging scenarios. Notably, the system identifies vulnerable pedestrian groups supporting more socially aware and inclusive robot behaviour.
title Vision-based Perception System for Automated Delivery Robot-Pedestrians Interactions
topic Robotics
Machine Learning
url https://arxiv.org/abs/2508.03541