Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Jingyu, Zhang, Chong, Liu, Fengqi, Fan, Ke, Zhou, Qianyu, Tan, Xin, Zhang, Zhizhong, Xie, Yuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908639390859264
author Gong, Jingyu
Zhang, Chong
Liu, Fengqi
Fan, Ke
Zhou, Qianyu
Tan, Xin
Zhang, Zhizhong
Xie, Yuan
author_facet Gong, Jingyu
Zhang, Chong
Liu, Fengqi
Fan, Ke
Zhou, Qianyu
Tan, Xin
Zhang, Zhizhong
Xie, Yuan
contents Scene-aware motion synthesis has been widely researched recently due to its numerous applications. Prevailing methods rely heavily on paired motion-scene data, while it is difficult to generalize to diverse scenes when trained only on a few specific ones. Thus, we propose a unified framework, termed Diffusion Implicit Policy (DIP), for scene-aware motion synthesis, where paired motion-scene data are no longer necessary. In this paper, we disentangle human-scene interaction from motion synthesis during training, and then introduce an interaction-based implicit policy into motion diffusion during inference. Synthesized motion can be derived through iterative diffusion denoising and implicit policy optimization, thus motion naturalness and interaction plausibility can be maintained simultaneously. For long-term motion synthesis, we introduce motion blending in joint rotation power space. The proposed method is evaluated on synthesized scenes with ShapeNet furniture, and real scenes from PROX and Replica. Results show that our framework presents better motion naturalness and interaction plausibility than cutting-edge methods. This also indicates the feasibility of utilizing the DIP for motion synthesis in more general tasks and versatile scenes. Code will be publicly available at https://github.com/jingyugong/DIP.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02261
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis
Gong, Jingyu
Zhang, Chong
Liu, Fengqi
Fan, Ke
Zhou, Qianyu
Tan, Xin
Zhang, Zhizhong
Xie, Yuan
Computer Vision and Pattern Recognition
Scene-aware motion synthesis has been widely researched recently due to its numerous applications. Prevailing methods rely heavily on paired motion-scene data, while it is difficult to generalize to diverse scenes when trained only on a few specific ones. Thus, we propose a unified framework, termed Diffusion Implicit Policy (DIP), for scene-aware motion synthesis, where paired motion-scene data are no longer necessary. In this paper, we disentangle human-scene interaction from motion synthesis during training, and then introduce an interaction-based implicit policy into motion diffusion during inference. Synthesized motion can be derived through iterative diffusion denoising and implicit policy optimization, thus motion naturalness and interaction plausibility can be maintained simultaneously. For long-term motion synthesis, we introduce motion blending in joint rotation power space. The proposed method is evaluated on synthesized scenes with ShapeNet furniture, and real scenes from PROX and Replica. Results show that our framework presents better motion naturalness and interaction plausibility than cutting-edge methods. This also indicates the feasibility of utilizing the DIP for motion synthesis in more general tasks and versatile scenes. Code will be publicly available at https://github.com/jingyugong/DIP.
title Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.02261