Autonomous Character-Scene Interaction Synthesis from Text Instruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Nan, He, Zimo, Wang, Zi, Li, Hongjie, Chen, Yixin, Huang, Siyuan, Zhu, Yixin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913537255800832
author Jiang, Nan
He, Zimo
Wang, Zi
Li, Hongjie
Chen, Yixin
Huang, Siyuan
Zhu, Yixin
author_facet Jiang, Nan
He, Zimo
Wang, Zi
Li, Hongjie
Chen, Yixin
Huang, Siyuan
Zhu, Yixin
contents Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These requirements pose challenges for current models, leading to a notable gap in automating the animation of characters from simple human inputs. This paper addresses this challenge by introducing a comprehensive framework for synthesizing multi-stage scene-aware interaction motions directly from a single text instruction and goal location. Our approach employs an auto-regressive diffusion model to synthesize the next motion segment, along with an autonomous scheduler predicting the transition for each action stage. To ensure that the synthesized motions are seamlessly integrated within the environment, we propose a scene representation that considers the local perception both at the start and the goal location. We further enhance the coherence of the generated motion by integrating frame embeddings with language input. Additionally, to support model training, we present a comprehensive motion-captured dataset comprising 16 hours of motion sequences in 120 indoor scenes covering 40 types of motions, each annotated with precise language descriptions. Experimental results demonstrate the efficacy of our method in generating high-quality, multi-stage motions closely aligned with environmental and textual conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03187
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Autonomous Character-Scene Interaction Synthesis from Text Instruction
Jiang, Nan
He, Zimo
Wang, Zi
Li, Hongjie
Chen, Yixin
Huang, Siyuan
Zhu, Yixin
Computer Vision and Pattern Recognition
Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage transitions. These requirements pose challenges for current models, leading to a notable gap in automating the animation of characters from simple human inputs. This paper addresses this challenge by introducing a comprehensive framework for synthesizing multi-stage scene-aware interaction motions directly from a single text instruction and goal location. Our approach employs an auto-regressive diffusion model to synthesize the next motion segment, along with an autonomous scheduler predicting the transition for each action stage. To ensure that the synthesized motions are seamlessly integrated within the environment, we propose a scene representation that considers the local perception both at the start and the goal location. We further enhance the coherence of the generated motion by integrating frame embeddings with language input. Additionally, to support model training, we present a comprehensive motion-captured dataset comprising 16 hours of motion sequences in 120 indoor scenes covering 40 types of motions, each annotated with precise language descriptions. Experimental results demonstrate the efficacy of our method in generating high-quality, multi-stage motions closely aligned with environmental and textual conditions.
title Autonomous Character-Scene Interaction Synthesis from Text Instruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.03187