Multi-Modal Grounded Planning and Efficient Replanning For Learning Embodied Agents with A Few Examples

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Taewoong, Kim, Byeonghwi, Choi, Jonghyun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910759608385536
author Kim, Taewoong
Kim, Byeonghwi
Choi, Jonghyun
author_facet Kim, Taewoong
Kim, Byeonghwi
Choi, Jonghyun
contents Learning a perception and reasoning module for robotic assistants to plan steps to perform complex tasks based on natural language instructions often requires large free-form language annotations, especially for short high-level instructions. To reduce the cost of annotation, large language models (LLMs) are used as a planner with few data. However, when elaborating the steps, even the state-of-the-art planner that uses LLMs mostly relies on linguistic common sense, often neglecting the status of the environment at command reception, resulting in inappropriate plans. To generate plans grounded in the environment, we propose FLARE (Few-shot Language with environmental Adaptive Replanning Embodied agent), which improves task planning using both language command and environmental perception. As language instructions often contain ambiguities or incorrect expressions, we additionally propose to correct the mistakes using visual cues from the agent. The proposed scheme allows us to use a few language pairs thanks to the visual cues and outperforms state-of-the-art approaches. Our code is available at https://github.com/snumprlab/flare.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17288
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Modal Grounded Planning and Efficient Replanning For Learning Embodied Agents with A Few Examples
Kim, Taewoong
Kim, Byeonghwi
Choi, Jonghyun
Robotics
Artificial Intelligence
Learning a perception and reasoning module for robotic assistants to plan steps to perform complex tasks based on natural language instructions often requires large free-form language annotations, especially for short high-level instructions. To reduce the cost of annotation, large language models (LLMs) are used as a planner with few data. However, when elaborating the steps, even the state-of-the-art planner that uses LLMs mostly relies on linguistic common sense, often neglecting the status of the environment at command reception, resulting in inappropriate plans. To generate plans grounded in the environment, we propose FLARE (Few-shot Language with environmental Adaptive Replanning Embodied agent), which improves task planning using both language command and environmental perception. As language instructions often contain ambiguities or incorrect expressions, we additionally propose to correct the mistakes using visual cues from the agent. The proposed scheme allows us to use a few language pairs thanks to the visual cues and outperforms state-of-the-art approaches. Our code is available at https://github.com/snumprlab/flare.
title Multi-Modal Grounded Planning and Efficient Replanning For Learning Embodied Agents with A Few Examples
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2412.17288