HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Xiaogang, Xie, Yiming, Wu, Zizhao, Jampani, Varun, Sun, Deqing, Jiang, Huaizu
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909676173524992
author Peng, Xiaogang
Xie, Yiming
Wu, Zizhao
Jampani, Varun
Sun, Deqing
Jiang, Huaizu
author_facet Peng, Xiaogang
Xie, Yiming
Wu, Zizhao
Jampani, Varun
Sun, Deqing
Jiang, Huaizu
contents We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch diffusion model (HOI-DM) to generate both human and object motions conditioned on the input text, and encourage coherent motions by a cross-attention communication module between the human and object motion generation branches. We also develop an affordance prediction diffusion model (APDM) to predict the contacting area between the human and object during the interactions driven by the textual prompt. The APDM is independent of the results by the HOI-DM and thus can correct potential errors by the latter. Moreover, it stochastically generates the contacting points to diversify the generated motions. Finally, we incorporate the estimated contacting points into the classifier-guidance to achieve accurate and close contact between humans and objects. To train and evaluate our approach, we annotate BEHAVE dataset with text descriptions. Experimental results on BEHAVE and OMOMO demonstrate that our approach produces realistic HOIs with various interactions and different types of objects.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06553
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
Peng, Xiaogang
Xie, Yiming
Wu, Zizhao
Jampani, Varun
Sun, Deqing
Jiang, Huaizu
Computer Vision and Pattern Recognition
We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch diffusion model (HOI-DM) to generate both human and object motions conditioned on the input text, and encourage coherent motions by a cross-attention communication module between the human and object motion generation branches. We also develop an affordance prediction diffusion model (APDM) to predict the contacting area between the human and object during the interactions driven by the textual prompt. The APDM is independent of the results by the HOI-DM and thus can correct potential errors by the latter. Moreover, it stochastically generates the contacting points to diversify the generated motions. Finally, we incorporate the estimated contacting points into the classifier-guidance to achieve accurate and close contact between humans and objects. To train and evaluate our approach, we annotate BEHAVE dataset with text descriptions. Experimental results on BEHAVE and OMOMO demonstrate that our approach produces realistic HOIs with various interactions and different types of objects.
title HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.06553