TANGO: Training-free Embodied AI Agents for Open-world Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ziliotto, Filippo, Campari, Tommaso, Serafini, Luciano, Ballan, Lamberto
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910744712314880
author Ziliotto, Filippo
Campari, Tommaso
Serafini, Luciano
Ballan, Lamberto
author_facet Ziliotto, Filippo
Campari, Tommaso
Serafini, Luciano
Ballan, Lamberto
contents Large Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. In this paper, we propose TANGO, an approach that extends the program composition via LLMs already observed for images, aiming to integrate those capabilities into embodied agents capable of observing and acting in the world. Specifically, by employing a simple PointGoal Navigation model combined with a memory-based exploration policy as a foundational primitive for guiding an agent through the world, we show how a single model can address diverse tasks without additional training. We task an LLM with composing the provided primitives to solve a specific task, using only a few in-context examples in the prompt. We evaluate our approach on three key Embodied AI tasks: Open-Set ObjectGoal Navigation, Multi-Modal Lifelong Navigation, and Open Embodied Question Answering, achieving state-of-the-art results without any specific fine-tuning in challenging zero-shot scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10402
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TANGO: Training-free Embodied AI Agents for Open-world Tasks
Ziliotto, Filippo
Campari, Tommaso
Serafini, Luciano
Ballan, Lamberto
Artificial Intelligence
Robotics
Large Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. In this paper, we propose TANGO, an approach that extends the program composition via LLMs already observed for images, aiming to integrate those capabilities into embodied agents capable of observing and acting in the world. Specifically, by employing a simple PointGoal Navigation model combined with a memory-based exploration policy as a foundational primitive for guiding an agent through the world, we show how a single model can address diverse tasks without additional training. We task an LLM with composing the provided primitives to solve a specific task, using only a few in-context examples in the prompt. We evaluate our approach on three key Embodied AI tasks: Open-Set ObjectGoal Navigation, Multi-Modal Lifelong Navigation, and Open Embodied Question Answering, achieving state-of-the-art results without any specific fine-tuning in challenging zero-shot scenarios.
title TANGO: Training-free Embodied AI Agents for Open-world Tasks
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2412.10402