GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saxena, Saumya, Buchanan, Blake, Paxton, Chris, Liu, Peiqi, Chen, Bingqing, Vaskevicius, Narunas, Palmieri, Luigi, Francis, Jonathan, Kroemer, Oliver
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914053147852800
author Saxena, Saumya
Buchanan, Blake
Paxton, Chris
Liu, Peiqi
Chen, Bingqing
Vaskevicius, Narunas
Palmieri, Luigi
Francis, Jonathan
Kroemer, Oliver
author_facet Saxena, Saumya
Buchanan, Blake
Paxton, Chris
Liu, Peiqi
Chen, Bingqing
Vaskevicius, Narunas
Palmieri, Luigi
Francis, Jonathan
Kroemer, Oliver
contents In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence. This problem remains challenging in robotics, due to the difficulties in obtaining useful semantic representations, updating these representations online, and leveraging prior world knowledge for efficient planning and exploration. To address these limitations, we propose GraphEQA, a novel approach that utilizes real-time 3D metric-semantic scene graphs (3DSGs) and task relevant images as multi-modal memory for grounding Vision-Language Models (VLMs) to perform EQA tasks in unseen environments. We employ a hierarchical planning approach that exploits the hierarchical nature of 3DSGs for structured planning and semantics-guided exploration. We evaluate GraphEQA in simulation on two benchmark datasets, HM-EQA and OpenEQA, and demonstrate that it outperforms key baselines by completing EQA tasks with higher success rates and fewer planning steps. We further demonstrate GraphEQA in multiple real-world home and office environments.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14480
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
Saxena, Saumya
Buchanan, Blake
Paxton, Chris
Liu, Peiqi
Chen, Bingqing
Vaskevicius, Narunas
Palmieri, Luigi
Francis, Jonathan
Kroemer, Oliver
Robotics
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence. This problem remains challenging in robotics, due to the difficulties in obtaining useful semantic representations, updating these representations online, and leveraging prior world knowledge for efficient planning and exploration. To address these limitations, we propose GraphEQA, a novel approach that utilizes real-time 3D metric-semantic scene graphs (3DSGs) and task relevant images as multi-modal memory for grounding Vision-Language Models (VLMs) to perform EQA tasks in unseen environments. We employ a hierarchical planning approach that exploits the hierarchical nature of 3DSGs for structured planning and semantics-guided exploration. We evaluate GraphEQA in simulation on two benchmark datasets, HM-EQA and OpenEQA, and demonstrate that it outperforms key baselines by completing EQA tasks with higher success rates and fewer planning steps. We further demonstrate GraphEQA in multiple real-world home and office environments.
title GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.14480