InterFusion: Text-Driven Generation of 3D Human-Object Interaction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dai, Sisi, Li, Wenhao, Sun, Haowen, Huang, Haibin, Ma, Chongyang, Huang, Hui, Xu, Kai, Hu, Ruizhen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914873023135744
author Dai, Sisi
Li, Wenhao
Sun, Haowen
Huang, Haibin
Ma, Chongyang
Huang, Hui
Xu, Kai
Hu, Ruizhen
author_facet Dai, Sisi
Li, Wenhao
Sun, Haowen
Huang, Haibin
Ma, Chongyang
Huang, Hui
Xu, Kai
Hu, Ruizhen
contents In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct text-to-3D methods in HOI, largely due to the lack of paired text-interaction data, and the inherent difficulties in simultaneously generating multiple concepts with complex spatial relationships. To effectively address these issues, we present InterFusion, a two-stage framework specifically designed for HOI generation. InterFusion involves human pose estimations derived from text as geometric priors, which simplifies the text-to-3D conversion process and introduces additional constraints for accurate object generation. At the first stage, InterFusion extracts 3D human poses from a synthesized image dataset depicting a wide range of interactions, subsequently mapping these poses to interaction descriptions. The second stage of InterFusion capitalizes on the latest developments in text-to-3D generation, enabling the production of realistic and high-quality 3D HOI scenes. This is achieved through a local-global optimization process, where the generation of human body and object is optimized separately, and jointly refined with a global optimization of the entire scene, ensuring a seamless and contextually coherent integration. Our experimental results affirm that InterFusion significantly outperforms existing state-of-the-art methods in 3D HOI generation.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15612
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InterFusion: Text-Driven Generation of 3D Human-Object Interaction
Dai, Sisi
Li, Wenhao
Sun, Haowen
Huang, Haibin
Ma, Chongyang
Huang, Hui
Xu, Kai
Hu, Ruizhen
Computer Vision and Pattern Recognition
In this study, we tackle the complex task of generating 3D human-object interactions (HOI) from textual descriptions in a zero-shot text-to-3D manner. We identify and address two key challenges: the unsatisfactory outcomes of direct text-to-3D methods in HOI, largely due to the lack of paired text-interaction data, and the inherent difficulties in simultaneously generating multiple concepts with complex spatial relationships. To effectively address these issues, we present InterFusion, a two-stage framework specifically designed for HOI generation. InterFusion involves human pose estimations derived from text as geometric priors, which simplifies the text-to-3D conversion process and introduces additional constraints for accurate object generation. At the first stage, InterFusion extracts 3D human poses from a synthesized image dataset depicting a wide range of interactions, subsequently mapping these poses to interaction descriptions. The second stage of InterFusion capitalizes on the latest developments in text-to-3D generation, enabling the production of realistic and high-quality 3D HOI scenes. This is achieved through a local-global optimization process, where the generation of human body and object is optimized separately, and jointly refined with a global optimization of the entire scene, ensuring a seamless and contextually coherent integration. Our experimental results affirm that InterFusion significantly outperforms existing state-of-the-art methods in 3D HOI generation.
title InterFusion: Text-Driven Generation of 3D Human-Object Interaction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.15612