Resolving Positional Ambiguity in Dialogues by Vision-Language Models for Robot Navigation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Kuan-Lin, Wei, Tzu-Ti, Yeh, Li-Tzu, Kao, Elaine, Tseng, Yu-Chee, Chen, Jen-Jee
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929547380785152
author Chen, Kuan-Lin
Wei, Tzu-Ti
Yeh, Li-Tzu
Kao, Elaine
Tseng, Yu-Chee
Chen, Jen-Jee
author_facet Chen, Kuan-Lin
Wei, Tzu-Ti
Yeh, Li-Tzu
Kao, Elaine
Tseng, Yu-Chee
Chen, Jen-Jee
contents We consider an autonomous navigation robot that can accept human commands through natural language to provide services in an indoor environment. These natural language commands may include time, position, object, and action components. However, we observe that the positional components within such commands usually refer to objects in the environment that may contain different levels of positional ambiguity. For example, the command "Go to the chair!" may be ambiguous when there are multiple chairs of the same type in a room. In order to disambiguate these commands, we employ a large language model and a large vision-language model to conduct multiple turns of conversations with the user. We propose a two-level approach that utilizes a vision-language model to map the meanings in natural language to a unique object ID in images and then performs another mapping from the unique object ID to a 3D depth map, thereby allowing the robot to navigate from its current position to the target position. To the best of our knowledge, this is the first work linking foundation models to the positional ambiguity issue.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12802
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Resolving Positional Ambiguity in Dialogues by Vision-Language Models for Robot Navigation
Chen, Kuan-Lin
Wei, Tzu-Ti
Yeh, Li-Tzu
Kao, Elaine
Tseng, Yu-Chee
Chen, Jen-Jee
Robotics
We consider an autonomous navigation robot that can accept human commands through natural language to provide services in an indoor environment. These natural language commands may include time, position, object, and action components. However, we observe that the positional components within such commands usually refer to objects in the environment that may contain different levels of positional ambiguity. For example, the command "Go to the chair!" may be ambiguous when there are multiple chairs of the same type in a room. In order to disambiguate these commands, we employ a large language model and a large vision-language model to conduct multiple turns of conversations with the user. We propose a two-level approach that utilizes a vision-language model to map the meanings in natural language to a unique object ID in images and then performs another mapping from the unique object ID to a 3D depth map, thereby allowing the robot to navigate from its current position to the target position. To the best of our knowledge, this is the first work linking foundation models to the positional ambiguity issue.
title Resolving Positional Ambiguity in Dialogues by Vision-Language Models for Robot Navigation
topic Robotics
url https://arxiv.org/abs/2410.12802