PAct: Part-Decomposed Single-View Articulated Object Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Qingming, Yao, Xinyue, Zhang, Shuyuan, Deng, Yueci, Liu, Guiliang, Liu, Zhen, Jia, Kui
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912907938234368
author Liu, Qingming
Yao, Xinyue
Zhang, Shuyuan
Deng, Yueci
Liu, Guiliang
Liu, Zhen
Jia, Kui
author_facet Liu, Qingming
Yao, Xinyue
Zhang, Shuyuan
Deng, Yueci
Liu, Guiliang
Liu, Zhen
Jia, Kui
contents Articulated objects are central to interactive 3D applications, including embodied AI, robotics, and VR/AR, where functional part decomposition and kinematic motion are essential. Yet producing high-fidelity articulated assets remains difficult to scale because it requires reliable part decomposition and kinematic rigging. Existing approaches largely fall into two paradigms: optimization-based reconstruction or distillation, which can be accurate but often takes tens of minutes to hours per instance, and inference-time methods that rely on template or part retrieval, producing plausible results that may not match the specific structure and appearance in the input observation. We introduce a part-centric generative framework for articulated object creation that synthesizes part geometry, composition, and articulation under explicit part-aware conditioning. Our representation models an object as a set of movable parts, each encoded by latent tokens augmented with part identity and articulation cues. Conditioned on a single image, the model generates articulated 3D assets that preserve instance-level correspondence while maintaining valid part structure and motion. The resulting approach avoids per-instance optimization, enables fast feed-forward inference, and supports controllable assembly and articulation, which are important for embodied interaction. Experiments on common articulated categories (e.g., drawers and doors) show improved input consistency, part accuracy, and articulation plausibility over optimization-based and retrieval-driven baselines, while substantially reducing inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14965
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PAct: Part-Decomposed Single-View Articulated Object Generation
Liu, Qingming
Yao, Xinyue
Zhang, Shuyuan
Deng, Yueci
Liu, Guiliang
Liu, Zhen
Jia, Kui
Computer Vision and Pattern Recognition
Robotics
Articulated objects are central to interactive 3D applications, including embodied AI, robotics, and VR/AR, where functional part decomposition and kinematic motion are essential. Yet producing high-fidelity articulated assets remains difficult to scale because it requires reliable part decomposition and kinematic rigging. Existing approaches largely fall into two paradigms: optimization-based reconstruction or distillation, which can be accurate but often takes tens of minutes to hours per instance, and inference-time methods that rely on template or part retrieval, producing plausible results that may not match the specific structure and appearance in the input observation. We introduce a part-centric generative framework for articulated object creation that synthesizes part geometry, composition, and articulation under explicit part-aware conditioning. Our representation models an object as a set of movable parts, each encoded by latent tokens augmented with part identity and articulation cues. Conditioned on a single image, the model generates articulated 3D assets that preserve instance-level correspondence while maintaining valid part structure and motion. The resulting approach avoids per-instance optimization, enables fast feed-forward inference, and supports controllable assembly and articulation, which are important for embodied interaction. Experiments on common articulated categories (e.g., drawers and doors) show improved input consistency, part accuracy, and articulation plausibility over optimization-based and retrieval-driven baselines, while substantially reducing inference time.
title PAct: Part-Decomposed Single-View Articulated Object Generation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2602.14965