DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Ruiqi, Wang, Xinjie, Liu, Liu, Guo, Chunle, Qiu, Jiaxiong, Li, Chongyi, Huang, Lichao, Su, Zhizhong, Cheng, Ming-Ming
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918037121138688
author Wu, Ruiqi
Wang, Xinjie
Liu, Liu
Guo, Chunle
Qiu, Jiaxiong
Li, Chongyi
Huang, Lichao
Su, Zhizhong
Cheng, Ming-Ming
author_facet Wu, Ruiqi
Wang, Xinjie
Liu, Liu
Guo, Chunle
Qiu, Jiaxiong
Li, Chongyi
Huang, Lichao
Su, Zhizhong
Cheng, Ming-Ming
contents We present DIPO, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for data collection, but at the same time provides important motion information, which is a reliable guide for predicting kinematic relationships between parts. Specifically, we propose a dual-image diffusion model that captures relationships between the image pair to generate part layouts and joint parameters. In addition, we introduce a Chain-of-Thought (CoT) based graph reasoner that explicitly infers part connectivity relationships. To further improve robustness and generalization on complex articulated objects, we develop a fully automated dataset expansion pipeline, name LEGO-Art, that enriches the diversity and complexity of PartNet-Mobility dataset. We propose PM-X, a large-scale dataset of complex articulated 3D objects, accompanied by rendered images, URDF annotations, and textual descriptions. Extensive experiments demonstrate that DIPO significantly outperforms existing baselines in both the resting state and the articulated state, while the proposed PM-X dataset further enhances generalization to diverse and structurally complex articulated objects. Our code and dataset will be released to the community upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20460
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
Wu, Ruiqi
Wang, Xinjie
Liu, Liu
Guo, Chunle
Qiu, Jiaxiong
Li, Chongyi
Huang, Lichao
Su, Zhizhong
Cheng, Ming-Ming
Computer Vision and Pattern Recognition
We present DIPO, a novel framework for the controllable generation of articulated 3D objects from a pair of images: one depicting the object in a resting state and the other in an articulated state. Compared to the single-image approach, our dual-image input imposes only a modest overhead for data collection, but at the same time provides important motion information, which is a reliable guide for predicting kinematic relationships between parts. Specifically, we propose a dual-image diffusion model that captures relationships between the image pair to generate part layouts and joint parameters. In addition, we introduce a Chain-of-Thought (CoT) based graph reasoner that explicitly infers part connectivity relationships. To further improve robustness and generalization on complex articulated objects, we develop a fully automated dataset expansion pipeline, name LEGO-Art, that enriches the diversity and complexity of PartNet-Mobility dataset. We propose PM-X, a large-scale dataset of complex articulated 3D objects, accompanied by rendered images, URDF annotations, and textual descriptions. Extensive experiments demonstrate that DIPO significantly outperforms existing baselines in both the resting state and the articulated state, while the proposed PM-X dataset further enhances generalization to diverse and structurally complex articulated objects. Our code and dataset will be released to the community upon publication.
title DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20460