JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Kaizhi, Zhou, Kaiwen, Gu, Jing, Fan, Yue, Wang, Jialu, Di, Zonglin, He, Xuehai, Wang, Xin Eric
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914017489977344
author Zheng, Kaizhi
Zhou, Kaiwen
Gu, Jing
Fan, Yue
Wang, Jialu
Di, Zonglin
He, Xuehai
Wang, Xin Eric
author_facet Zheng, Kaizhi
Zhou, Kaiwen
Gu, Jing
Fan, Yue
Wang, Jialu
Di, Zonglin
He, Xuehai
Wang, Xin Eric
contents Building a conversational embodied agent to execute real-life tasks has been a long-standing yet quite challenging research goal, as it requires effective human-agent communication, multi-modal understanding, long-range sequential decision making, etc. Traditional symbolic methods have scaling and generalization issues, while end-to-end deep learning models suffer from data scarcity and high task complexity, and are often hard to explain. To benefit from both worlds, we propose JARVIS, a neuro-symbolic commonsense reasoning framework for modular, generalizable, and interpretable conversational embodied agents. First, it acquires symbolic representations by prompting large language models (LLMs) for language understanding and sub-goal planning, and by constructing semantic maps from visual observations. Then the symbolic module reasons for sub-goal planning and action generation based on task- and action-level common sense. Extensive experiments on the TEACh dataset validate the efficacy and efficiency of our JARVIS framework, which achieves state-of-the-art (SOTA) results on all three dialog-based embodied tasks, including Execution from Dialog History (EDH), Trajectory from Dialog (TfD), and Two-Agent Task Completion (TATC) (e.g., our method boosts the unseen Success Rate on EDH from 6.1\% to 15.8\%). Moreover, we systematically analyze the essential factors that affect the task performance and also demonstrate the superiority of our method in few-shot settings. Our JARVIS model ranks first in the Alexa Prize SimBot Public Benchmark Challenge.
format Preprint
id arxiv_https___arxiv_org_abs_2208_13266
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
Zheng, Kaizhi
Zhou, Kaiwen
Gu, Jing
Fan, Yue
Wang, Jialu
Di, Zonglin
He, Xuehai
Wang, Xin Eric
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
Building a conversational embodied agent to execute real-life tasks has been a long-standing yet quite challenging research goal, as it requires effective human-agent communication, multi-modal understanding, long-range sequential decision making, etc. Traditional symbolic methods have scaling and generalization issues, while end-to-end deep learning models suffer from data scarcity and high task complexity, and are often hard to explain. To benefit from both worlds, we propose JARVIS, a neuro-symbolic commonsense reasoning framework for modular, generalizable, and interpretable conversational embodied agents. First, it acquires symbolic representations by prompting large language models (LLMs) for language understanding and sub-goal planning, and by constructing semantic maps from visual observations. Then the symbolic module reasons for sub-goal planning and action generation based on task- and action-level common sense. Extensive experiments on the TEACh dataset validate the efficacy and efficiency of our JARVIS framework, which achieves state-of-the-art (SOTA) results on all three dialog-based embodied tasks, including Execution from Dialog History (EDH), Trajectory from Dialog (TfD), and Two-Agent Task Completion (TATC) (e.g., our method boosts the unseen Success Rate on EDH from 6.1\% to 15.8\%). Moreover, we systematically analyze the essential factors that affect the task performance and also demonstrate the superiority of our method in few-shot settings. Our JARVIS model ranks first in the Alexa Prize SimBot Public Benchmark Challenge.
title JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2208.13266