Controllable Human-Object Interaction Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiaman, Clegg, Alexander, Mottaghi, Roozbeh, Wu, Jiajun, Puig, Xavier, Liu, C. Karen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914868429324288
author Li, Jiaman
Clegg, Alexander
Mottaghi, Roozbeh
Wu, Jiajun
Puig, Xavier
Liu, C. Karen
author_facet Li, Jiaman
Clegg, Alexander
Mottaghi, Roozbeh
Wu, Jiajun
Puig, Xavier
Liu, C. Karen
contents Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human motion guided by language descriptions in 3D scenes. We propose Controllable Human-Object Interaction Synthesis (CHOIS), an approach that generates object motion and human motion simultaneously using a conditional diffusion model given a language description, initial object and human states, and sparse object waypoints. Here, language descriptions inform style and intent, and waypoints, which can be effectively extracted from high-level planning, ground the motion in the scene. Naively applying a diffusion model fails to predict object motion aligned with the input waypoints; it also cannot ensure the realism of interactions that require precise hand-object and human-floor contact. To overcome these problems, we introduce an object geometry loss as additional supervision to improve the matching between generated object motion and input object waypoints; we also design guidance terms to enforce contact constraints during the sampling process of the trained diffusion model. We demonstrate that our learned interaction module can synthesize realistic human-object interactions, adhering to provided textual descriptions and sparse waypoint conditions. Additionally, our module seamlessly integrates with a path planning module, enabling the generation of long-term interactions in 3D environments.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03913
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Controllable Human-Object Interaction Synthesis
Li, Jiaman
Clegg, Alexander
Mottaghi, Roozbeh
Wu, Jiajun
Puig, Xavier
Liu, C. Karen
Computer Vision and Pattern Recognition
Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human motion guided by language descriptions in 3D scenes. We propose Controllable Human-Object Interaction Synthesis (CHOIS), an approach that generates object motion and human motion simultaneously using a conditional diffusion model given a language description, initial object and human states, and sparse object waypoints. Here, language descriptions inform style and intent, and waypoints, which can be effectively extracted from high-level planning, ground the motion in the scene. Naively applying a diffusion model fails to predict object motion aligned with the input waypoints; it also cannot ensure the realism of interactions that require precise hand-object and human-floor contact. To overcome these problems, we introduce an object geometry loss as additional supervision to improve the matching between generated object motion and input object waypoints; we also design guidance terms to enforce contact constraints during the sampling process of the trained diffusion model. We demonstrate that our learned interaction module can synthesize realistic human-object interactions, adhering to provided textual descriptions and sparse waypoint conditions. Additionally, our module seamlessly integrates with a path planning module, enabling the generation of long-term interactions in 3D environments.
title Controllable Human-Object Interaction Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.03913