Map-based Modular Approach for Zero-shot Embodied Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakamoto, Koya, Azuma, Daichi, Miyanishi, Taiki, Kurita, Shuhei, Kawanabe, Motoaki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909347205873664
author Sakamoto, Koya
Azuma, Daichi
Miyanishi, Taiki
Kurita, Shuhei
Kawanabe, Motoaki
author_facet Sakamoto, Koya
Azuma, Daichi
Miyanishi, Taiki
Kurita, Shuhei
Kawanabe, Motoaki
contents Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This paper presents a map-based modular approach to EQA, enabling real-world robots to explore and map unknown environments. By leveraging foundation models, our method facilitates answering a diverse range of questions using natural language. We conducted extensive experiments in both virtual and real-world settings, demonstrating the robustness of our approach in navigating and comprehending queries within unknown environments.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16559
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Map-based Modular Approach for Zero-shot Embodied Question Answering
Sakamoto, Koya
Azuma, Daichi
Miyanishi, Taiki
Kurita, Shuhei
Kawanabe, Motoaki
Robotics
Computer Vision and Pattern Recognition
Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods often rely on simulated environments and operate with limited vocabularies. This paper presents a map-based modular approach to EQA, enabling real-world robots to explore and map unknown environments. By leveraging foundation models, our method facilitates answering a diverse range of questions using natural language. We conducted extensive experiments in both virtual and real-world settings, demonstrating the robustness of our approach in navigating and comprehending queries within unknown environments.
title Map-based Modular Approach for Zero-shot Embodied Question Answering
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.16559