GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Hao, Fan, Lei, Di, Donglin, Liu, Shaohui
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912745995108352
author Sun, Hao
Fan, Lei
Di, Donglin
Liu, Shaohui
author_facet Sun, Hao
Fan, Lei
Di, Donglin
Liu, Shaohui
contents Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process. First, we fine-tune a point cloud generation model to produce a coarse representation of objects from text prompts. Given the inherent connection between articulated objects and graph structures, we design a hypergraph-based learning method to refine these coarse representations, representing object parts as graph vertices. Finally, leveraging a diffusion model, the joints of articulated objects-represented as graph edges-are generated based on the object parts. Extensive qualitative and quantitative experiments on the PartNet-Mobility dataset demonstrate the effectiveness of our approach, achieving superior performance over previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03566
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
Sun, Hao
Fan, Lei
Di, Donglin
Liu, Shaohui
Computer Vision and Pattern Recognition
Multimedia
Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process. First, we fine-tune a point cloud generation model to produce a coarse representation of objects from text prompts. Given the inherent connection between articulated objects and graph structures, we design a hypergraph-based learning method to refine these coarse representations, representing object parts as graph vertices. Finally, leveraging a diffusion model, the joints of articulated objects-represented as graph edges-are generated based on the object parts. Extensive qualitative and quantitative experiments on the PartNet-Mobility dataset demonstrate the effectiveness of our approach, achieving superior performance over previous methods.
title GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2512.03566