Pointing-Guided Target Estimation via Transformer-Based Attention

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Müller, Luca, Ali, Hassan, Allgeuer, Philipp, Gajdošech, Lukáš, Wermter, Stefan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908521274015744
author Müller, Luca
Ali, Hassan
Allgeuer, Philipp
Gajdošech, Lukáš
Wermter, Stefan
author_facet Müller, Luca
Ali, Hassan
Allgeuer, Philipp
Gajdošech, Lukáš
Wermter, Stefan
contents Deictic gestures, like pointing, are a fundamental form of non-verbal communication, enabling humans to direct attention to specific objects or locations. This capability is essential in Human-Robot Interaction (HRI), where robots should be able to predict human intent and anticipate appropriate responses. In this work, we propose the Multi-Modality Inter-TransFormer (MM-ITF), a modular architecture to predict objects in a controlled tabletop scenario with the NICOL robot, where humans indicate targets through natural pointing gestures. Leveraging inter-modality attention, MM-ITF maps 2D pointing gestures to object locations, assigns a likelihood score to each, and identifies the most likely target. Our results demonstrate that the method can accurately predict the intended object using monocular RGB data, thus enabling intuitive and accessible human-robot collaboration. To evaluate the performance, we introduce a patch confusion matrix, providing insights into the model's predictions across candidate object locations. Code available at: https://github.com/lucamuellercode/MMITF.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pointing-Guided Target Estimation via Transformer-Based Attention
Müller, Luca
Ali, Hassan
Allgeuer, Philipp
Gajdošech, Lukáš
Wermter, Stefan
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.9; I.2.10; I.2.6
Deictic gestures, like pointing, are a fundamental form of non-verbal communication, enabling humans to direct attention to specific objects or locations. This capability is essential in Human-Robot Interaction (HRI), where robots should be able to predict human intent and anticipate appropriate responses. In this work, we propose the Multi-Modality Inter-TransFormer (MM-ITF), a modular architecture to predict objects in a controlled tabletop scenario with the NICOL robot, where humans indicate targets through natural pointing gestures. Leveraging inter-modality attention, MM-ITF maps 2D pointing gestures to object locations, assigns a likelihood score to each, and identifies the most likely target. Our results demonstrate that the method can accurately predict the intended object using monocular RGB data, thus enabling intuitive and accessible human-robot collaboration. To evaluate the performance, we introduce a patch confusion matrix, providing insights into the model's predictions across candidate object locations. Code available at: https://github.com/lucamuellercode/MMITF.
title Pointing-Guided Target Estimation via Transformer-Based Attention
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.9; I.2.10; I.2.6
url https://arxiv.org/abs/2509.05031