A Multimodal Depth-Aware Method For Embodied Reference Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917446237028352 |
|---|---|
| author | Eyiokur, Fevziye Irem Yaman, Dogucan Ekenel, Hazım Kemal Waibel, Alexander |
| author_facet | Eyiokur, Fevziye Irem Yaman, Dogucan Ekenel, Hazım Kemal Waibel, Alexander |
| contents | Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08278 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Multimodal Depth-Aware Method For Embodied Reference Understanding Eyiokur, Fevziye Irem Yaman, Dogucan Ekenel, Hazım Kemal Waibel, Alexander Computer Vision and Pattern Recognition Human-Computer Interaction Robotics Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in the scene. To address these challenges, we propose a novel ERU framework that jointly leverages LLM-based data augmentation, depth-map modality, and a depth-aware decision module. This design enables robust integration of linguistic and embodied cues, improving disambiguation in complex or cluttered environments. Experimental results on two datasets demonstrate that our approach significantly outperforms existing baselines, achieving more accurate and reliable referent detection. |
| title | A Multimodal Depth-Aware Method For Embodied Reference Understanding |
| topic | Computer Vision and Pattern Recognition Human-Computer Interaction Robotics |
| url | https://arxiv.org/abs/2510.08278 |