A Multimodal Depth-Aware Method For Embodied Reference Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Eyiokur, Fevziye Irem, Yaman, Dogucan, Ekenel, Hazım Kemal, Waibel, Alexander
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917446237028352
author Eyiokur, Fevziye Irem
Yaman, Dogucan
Ekenel, Hazım Kemal
Waibel, Alexander
author_facet Eyiokur, Fevziye Irem
Yaman, Dogucan
Ekenel, Hazım Kemal
Waibel, Alexander
contents Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multimodal Depth-Aware Method For Embodied Reference Understanding
Eyiokur, Fevziye Irem
Yaman, Dogucan
Ekenel, Hazım Kemal
Waibel, Alexander
Computer Vision and Pattern Recognition
Human-Computer Interaction
Robotics
Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection.
title A Multimodal Depth-Aware Method For Embodied Reference Understanding
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
Robotics
url https://arxiv.org/abs/2510.08278