PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chang, Matthew, Chhablani, Gunjan, Clegg, Alexander, Cote, Mikael Dallaire, Desai, Ruta, Hlavac, Michal, Karashchuk, Vladimir, Krantz, Jacob, Mottaghi, Roozbeh, Parashar, Priyam, Patki, Siddharth, Prasad, Ishita, Puig, Xavier, Rai, Akshara, Ramrakhya, Ram, Tran, Daniel, Truong, Joanne, Turner, John M., Undersander, Eric, Yang, Tsung-Yen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909374351409152
author Chang, Matthew
Chhablani, Gunjan
Clegg, Alexander
Cote, Mikael Dallaire
Desai, Ruta
Hlavac, Michal
Karashchuk, Vladimir
Krantz, Jacob
Mottaghi, Roozbeh
Parashar, Priyam
Patki, Siddharth
Prasad, Ishita
Puig, Xavier
Rai, Akshara
Ramrakhya, Ram
Tran, Daniel
Truong, Joanne
Turner, John M.
Undersander, Eric
Yang, Tsung-Yen
author_facet Chang, Matthew
Chhablani, Gunjan
Clegg, Alexander
Cote, Mikael Dallaire
Desai, Ruta
Hlavac, Michal
Karashchuk, Vladimir
Krantz, Jacob
Mottaghi, Roozbeh
Parashar, Priyam
Patki, Siddharth
Prasad, Ishita
Puig, Xavier
Rai, Akshara
Ramrakhya, Ram
Tran, Daniel
Truong, Joanne
Turner, John M.
Undersander, Eric
Yang, Tsung-Yen
contents We present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhibit characteristics of everyday tasks, such as spatial, temporal, and heterogeneous agent capability constraints. We employ a semi-automated task generation pipeline using Large Language Models (LLMs), incorporating simulation in the loop for grounding and verification. PARTNR stands as the largest benchmark of its kind, comprising 100,000 natural language tasks, spanning 60 houses and 5,819 unique objects. We analyze state-of-the-art LLMs on PARTNR tasks, across the axes of planning, perception and skill execution. The analysis reveals significant limitations in SoTA models, such as poor coordination and failures in task tracking and recovery from errors. When LLMs are paired with real humans, they require 1.5x as many steps as two humans collaborating and 1.1x more steps than a single human, underscoring the potential for improvement in these models. We further show that fine-tuning smaller LLMs with planning data can achieve performance on par with models 9 times larger, while being 8.6x faster at inference. Overall, PARTNR highlights significant challenges facing collaborative embodied agents and aims to drive research in this direction.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
Chang, Matthew
Chhablani, Gunjan
Clegg, Alexander
Cote, Mikael Dallaire
Desai, Ruta
Hlavac, Michal
Karashchuk, Vladimir
Krantz, Jacob
Mottaghi, Roozbeh
Parashar, Priyam
Patki, Siddharth
Prasad, Ishita
Puig, Xavier
Rai, Akshara
Ramrakhya, Ram
Tran, Daniel
Truong, Joanne
Turner, John M.
Undersander, Eric
Yang, Tsung-Yen
Robotics
Artificial Intelligence
We present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhibit characteristics of everyday tasks, such as spatial, temporal, and heterogeneous agent capability constraints. We employ a semi-automated task generation pipeline using Large Language Models (LLMs), incorporating simulation in the loop for grounding and verification. PARTNR stands as the largest benchmark of its kind, comprising 100,000 natural language tasks, spanning 60 houses and 5,819 unique objects. We analyze state-of-the-art LLMs on PARTNR tasks, across the axes of planning, perception and skill execution. The analysis reveals significant limitations in SoTA models, such as poor coordination and failures in task tracking and recovery from errors. When LLMs are paired with real humans, they require 1.5x as many steps as two humans collaborating and 1.1x more steps than a single human, underscoring the potential for improvement in these models. We further show that fine-tuning smaller LLMs with planning data can achieve performance on par with models 9 times larger, while being 8.6x faster at inference. Overall, PARTNR highlights significant challenges facing collaborative embodied agents and aims to drive research in this direction.
title PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2411.00081