MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hong, Yining, Zheng, Zishuo, Chen, Peihao, Wang, Yian, Li, Junyan, Gan, Chuang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909075036438528
author Hong, Yining
Zheng, Zishuo
Chen, Peihao
Wang, Yian
Li, Junyan
Gan, Chuang
author_facet Hong, Yining
Zheng, Zishuo
Chen, Peihao
Wang, Yian
Li, Junyan
Gan, Chuang
contents Human beings possess the capability to multiply a melange of multisensory cues while actively exploring and interacting with the 3D world. Current multi-modal large language models, however, passively absorb sensory data as inputs, lacking the capacity to actively interact with the objects in the 3D environment and dynamically collect their multisensory information. To usher in the study of this area, we propose MultiPLY, a multisensory embodied large language model that could incorporate multisensory interactive data, including visual, audio, tactile, and thermal information into large language models, thereby establishing the correlation among words, actions, and percepts. To this end, we first collect Multisensory Universe, a large-scale multisensory interaction dataset comprising 500k data by deploying an LLM-powered embodied agent to engage with the 3D environment. To perform instruction tuning with pre-trained LLM on such generated data, we first encode the 3D scene as abstracted object-centric representations and then introduce action tokens denoting that the embodied agent takes certain actions within the environment, as well as state tokens that represent the multisensory state observations of the agent at each time step. In the inference time, MultiPLY could generate action tokens, instructing the agent to take the action in the environment and obtain the next multisensory state observation. The observation is then appended back to the LLM via state tokens to generate subsequent text or action tokens. We demonstrate that MultiPLY outperforms baselines by a large margin through a diverse set of embodied tasks involving object retrieval, tool use, multisensory captioning, and task decomposition.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08577
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
Hong, Yining
Zheng, Zishuo
Chen, Peihao
Wang, Yian
Li, Junyan
Gan, Chuang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
Human beings possess the capability to multiply a melange of multisensory cues while actively exploring and interacting with the 3D world. Current multi-modal large language models, however, passively absorb sensory data as inputs, lacking the capacity to actively interact with the objects in the 3D environment and dynamically collect their multisensory information. To usher in the study of this area, we propose MultiPLY, a multisensory embodied large language model that could incorporate multisensory interactive data, including visual, audio, tactile, and thermal information into large language models, thereby establishing the correlation among words, actions, and percepts. To this end, we first collect Multisensory Universe, a large-scale multisensory interaction dataset comprising 500k data by deploying an LLM-powered embodied agent to engage with the 3D environment. To perform instruction tuning with pre-trained LLM on such generated data, we first encode the 3D scene as abstracted object-centric representations and then introduce action tokens denoting that the embodied agent takes certain actions within the environment, as well as state tokens that represent the multisensory state observations of the agent at each time step. In the inference time, MultiPLY could generate action tokens, instructing the agent to take the action in the environment and obtain the next multisensory state observation. The observation is then appended back to the LLM via state tokens to generate subsequent text or action tokens. We demonstrate that MultiPLY outperforms baselines by a large margin through a diverse set of embodied tasks involving object retrieval, tool use, multisensory captioning, and task decomposition.
title MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
url https://arxiv.org/abs/2401.08577