ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929625601409024 |
|---|---|
| author | Shen, Ying Bis, Daniel Lu, Cynthia Lourentzou, Ismini |
| author_facet | Shen, Ying Bis, Daniel Lu, Cynthia Lourentzou, Ismini |
| contents | The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language benchmarks, most prior work has focused on building agents that follow instructions rather than endowing agents the ability to ask questions to actively resolve ambiguities arising naturally in embodied environments. To address this gap, we propose an Embodied Learning-By-Asking (ELBA) model that learns when and what questions to ask to dynamically acquire additional information for completing the task. We evaluate ELBA on the TEACh vision-dialog navigation and task completion dataset. Experimental results show that the proposed method achieves improved task performance compared to baseline models without question-answering capabilities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2302_04865 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion Shen, Ying Bis, Daniel Lu, Cynthia Lourentzou, Ismini Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language benchmarks, most prior work has focused on building agents that follow instructions rather than endowing agents the ability to ask questions to actively resolve ambiguities arising naturally in embodied environments. To address this gap, we propose an Embodied Learning-By-Asking (ELBA) model that learns when and what questions to ask to dynamically acquire additional information for completing the task. We evaluate ELBA on the TEACh vision-dialog navigation and task completion dataset. Experimental results show that the proposed method achieves improved task performance compared to baseline models without question-answering capabilities. |
| title | ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2302.04865 |