GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914053147852800 |
|---|---|
| author | Saxena, Saumya Buchanan, Blake Paxton, Chris Liu, Peiqi Chen, Bingqing Vaskevicius, Narunas Palmieri, Luigi Francis, Jonathan Kroemer, Oliver |
| author_facet | Saxena, Saumya Buchanan, Blake Paxton, Chris Liu, Peiqi Chen, Bingqing Vaskevicius, Narunas Palmieri, Luigi Francis, Jonathan Kroemer, Oliver |
| contents | In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence. This problem remains challenging in robotics, due to the difficulties in obtaining useful semantic representations, updating these representations online, and leveraging prior world knowledge for efficient planning and exploration. To address these limitations, we propose GraphEQA, a novel approach that utilizes real-time 3D metric-semantic scene graphs (3DSGs) and task relevant images as multi-modal memory for grounding Vision-Language Models (VLMs) to perform EQA tasks in unseen environments. We employ a hierarchical planning approach that exploits the hierarchical nature of 3DSGs for structured planning and semantics-guided exploration. We evaluate GraphEQA in simulation on two benchmark datasets, HM-EQA and OpenEQA, and demonstrate that it outperforms key baselines by completing EQA tasks with higher success rates and fewer planning steps. We further demonstrate GraphEQA in multiple real-world home and office environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_14480 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering Saxena, Saumya Buchanan, Blake Paxton, Chris Liu, Peiqi Chen, Bingqing Vaskevicius, Narunas Palmieri, Luigi Francis, Jonathan Kroemer, Oliver Robotics Computation and Language Computer Vision and Pattern Recognition Machine Learning In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment to answer a situated question with confidence. This problem remains challenging in robotics, due to the difficulties in obtaining useful semantic representations, updating these representations online, and leveraging prior world knowledge for efficient planning and exploration. To address these limitations, we propose GraphEQA, a novel approach that utilizes real-time 3D metric-semantic scene graphs (3DSGs) and task relevant images as multi-modal memory for grounding Vision-Language Models (VLMs) to perform EQA tasks in unseen environments. We employ a hierarchical planning approach that exploits the hierarchical nature of 3DSGs for structured planning and semantics-guided exploration. We evaluate GraphEQA in simulation on two benchmark datasets, HM-EQA and OpenEQA, and demonstrate that it outperforms key baselines by completing EQA tasks with higher success rates and fewer planning steps. We further demonstrate GraphEQA in multiple real-world home and office environments. |
| title | GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering |
| topic | Robotics Computation and Language Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2412.14480 |