Multimodal LLM Guided Exploration and Active Mapping using Fisher Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Wen, Lei, Boshu, Ashton, Katrina, Daniilidis, Kostas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908519263895552
author Jiang, Wen
Lei, Boshu
Ashton, Katrina
Daniilidis, Kostas
author_facet Jiang, Wen
Lei, Boshu
Ashton, Katrina
Daniilidis, Kostas
contents We present an active mapping system that plans for both long-horizon exploration goals and short-term actions using a 3D Gaussian Splatting (3DGS) representation. Existing methods either do not take advantage of recent developments in multimodal Large Language Models (LLM) or do not consider challenges in localization uncertainty, which is critical in embodied agents. We propose employing multimodal LLMs for long-horizon planning in conjunction with detailed motion planning using our information-based objective. By leveraging high-quality view synthesis from our 3DGS representation, our method employs a multimodal LLM as a zero-shot planner for long-horizon exploration goals from the semantic perspective. We also introduce an uncertainty-aware path proposal and selection algorithm that balances the dual objectives of maximizing the information gain for the environment while minimizing the cost of localization errors. Experiments conducted on the Gibson and Habitat-Matterport 3D datasets demonstrate state-of-the-art results of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17422
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal LLM Guided Exploration and Active Mapping using Fisher Information
Jiang, Wen
Lei, Boshu
Ashton, Katrina
Daniilidis, Kostas
Robotics
Computer Vision and Pattern Recognition
We present an active mapping system that plans for both long-horizon exploration goals and short-term actions using a 3D Gaussian Splatting (3DGS) representation. Existing methods either do not take advantage of recent developments in multimodal Large Language Models (LLM) or do not consider challenges in localization uncertainty, which is critical in embodied agents. We propose employing multimodal LLMs for long-horizon planning in conjunction with detailed motion planning using our information-based objective. By leveraging high-quality view synthesis from our 3DGS representation, our method employs a multimodal LLM as a zero-shot planner for long-horizon exploration goals from the semantic perspective. We also introduce an uncertainty-aware path proposal and selection algorithm that balances the dual objectives of maximizing the information gain for the environment while minimizing the cost of localization errors. Experiments conducted on the Gibson and Habitat-Matterport 3D datasets demonstrate state-of-the-art results of the proposed method.
title Multimodal LLM Guided Exploration and Active Mapping using Fisher Information
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.17422