ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Ying, Bis, Daniel, Lu, Cynthia, Lourentzou, Ismini
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929625601409024
author Shen, Ying
Bis, Daniel
Lu, Cynthia
Lourentzou, Ismini
author_facet Shen, Ying
Bis, Daniel
Lu, Cynthia
Lourentzou, Ismini
contents The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language benchmarks, most prior work has focused on building agents that follow instructions rather than endowing agents the ability to ask questions to actively resolve ambiguities arising naturally in embodied environments. To address this gap, we propose an Embodied Learning-By-Asking (ELBA) model that learns when and what questions to ask to dynamically acquire additional information for completing the task. We evaluate ELBA on the TEACh vision-dialog navigation and task completion dataset. Experimental results show that the proposed method achieves improved task performance compared to baseline models without question-answering capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2302_04865
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
Shen, Ying
Bis, Daniel
Lu, Cynthia
Lourentzou, Ismini
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language benchmarks, most prior work has focused on building agents that follow instructions rather than endowing agents the ability to ask questions to actively resolve ambiguities arising naturally in embodied environments. To address this gap, we propose an Embodied Learning-By-Asking (ELBA) model that learns when and what questions to ask to dynamically acquire additional information for completing the task. We evaluate ELBA on the TEACh vision-dialog navigation and task completion dataset. Experimental results show that the proposed method achieves improved task performance compared to baseline models without question-answering capabilities.
title ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2302.04865