SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Yukai, Li, Weiyu, Wang, Zihao, Li, Hongyang, Chen, Xingyu, Tan, Ping, Zhang, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911314513756160
author Shi, Yukai
Li, Weiyu
Wang, Zihao
Li, Hongyang
Chen, Xingyu
Tan, Ping
Zhang, Lei
author_facet Shi, Yukai
Li, Weiyu
Wang, Zihao
Li, Hongyang
Chen, Xingyu
Tan, Ping
Zhang, Lei
contents We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry and accurate poses under severe occlusion and open-set settings. To address these issues, we first decouple the de-occlusion model from 3D object generation, and enhance it by leveraging image datasets and collected de-occlusion datasets for much more diverse open-set occlusion patterns. Then, we propose a unified pose estimation model that integrates global and local mechanisms for both self-attention and cross-attention to improve accuracy. Besides, we construct an open-set 3D scene dataset to further extend the generalization of the pose estimation model. Comprehensive experiments demonstrate the superiority of our decoupled framework on both indoor and open-set scenes. Our codes and datasets is released at https://idea-research.github.io/SceneMaker/.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10957
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
Shi, Yukai
Li, Weiyu
Wang, Zihao
Li, Hongyang
Chen, Xingyu
Tan, Ping
Zhang, Lei
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry and accurate poses under severe occlusion and open-set settings. To address these issues, we first decouple the de-occlusion model from 3D object generation, and enhance it by leveraging image datasets and collected de-occlusion datasets for much more diverse open-set occlusion patterns. Then, we propose a unified pose estimation model that integrates global and local mechanisms for both self-attention and cross-attention to improve accuracy. Besides, we construct an open-set 3D scene dataset to further extend the generalization of the pose estimation model. Comprehensive experiments demonstrate the superiority of our decoupled framework on both indoor and open-set scenes. Our codes and datasets is released at https://idea-research.github.io/SceneMaker/.
title SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.10957