MotionChain: Conversational Motion Controllers via Multimodal Prompts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Biao, Chen, Xin, Zhang, Chi, Yin, Fukun, Li, Zhuoyuan, YU, Gang, Fan, Jiayuan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911825540415488
author Jiang, Biao
Chen, Xin
Zhang, Chi
Yin, Fukun
Li, Zhuoyuan
YU, Gang
Fan, Jiayuan
author_facet Jiang, Biao
Chen, Xin
Zhang, Chi
Yin, Fukun
Li, Zhuoyuan
YU, Gang
Fan, Jiayuan
contents Recent advancements in language models have demonstrated their adeptness in conducting multi-turn dialogues and retaining conversational context. However, this proficiency remains largely unexplored in other multimodal generative models, particularly in human motion models. By integrating multi-turn conversations in controlling continuous virtual human movements, generative human motion models can achieve an intuitive and step-by-step process of human task execution for humanoid robotics, game agents, or other embodied systems. In this work, we present MotionChain, a conversational human motion controller to generate continuous and long-term human motion through multimodal prompts. Specifically, MotionChain consists of multi-modal tokenizers that transform various data types such as text, image, and motion, into discrete tokens, coupled with a Vision-Motion-aware Language model. By leveraging large-scale language, vision-language, and vision-motion data to assist motion-related generation tasks, MotionChain thus comprehends each instruction in multi-turn conversation and generates human motions followed by these prompts. Extensive experiments validate the efficacy of MotionChain, demonstrating state-of-the-art performance in conversational motion generation, as well as more intuitive manners of controlling and interacting with virtual humans.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01700
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MotionChain: Conversational Motion Controllers via Multimodal Prompts
Jiang, Biao
Chen, Xin
Zhang, Chi
Yin, Fukun
Li, Zhuoyuan
YU, Gang
Fan, Jiayuan
Computer Vision and Pattern Recognition
Recent advancements in language models have demonstrated their adeptness in conducting multi-turn dialogues and retaining conversational context. However, this proficiency remains largely unexplored in other multimodal generative models, particularly in human motion models. By integrating multi-turn conversations in controlling continuous virtual human movements, generative human motion models can achieve an intuitive and step-by-step process of human task execution for humanoid robotics, game agents, or other embodied systems. In this work, we present MotionChain, a conversational human motion controller to generate continuous and long-term human motion through multimodal prompts. Specifically, MotionChain consists of multi-modal tokenizers that transform various data types such as text, image, and motion, into discrete tokens, coupled with a Vision-Motion-aware Language model. By leveraging large-scale language, vision-language, and vision-motion data to assist motion-related generation tasks, MotionChain thus comprehends each instruction in multi-turn conversation and generates human motions followed by these prompts. Extensive experiments validate the efficacy of MotionChain, demonstrating state-of-the-art performance in conversational motion generation, as well as more intuitive manners of controlling and interacting with virtual humans.
title MotionChain: Conversational Motion Controllers via Multimodal Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.01700