PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xiang, Yoo, Sohyun, Wu, Hongrui, Li, Chuan, Xie, Jianwen, Tu, Zhuowen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918375405387776
author Zhang, Xiang
Yoo, Sohyun
Wu, Hongrui
Li, Chuan
Xie, Jianwen
Tu, Zhuowen
author_facet Zhang, Xiang
Yoo, Sohyun
Wu, Hongrui
Li, Chuan
Xie, Jianwen
Tu, Zhuowen
contents We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist-ready meshes in a single forward pass. Building on recent advances in mesh generative models, we augment a point-cloud encoder with pixel-aligned image features and global scene context via cross-attention, enabling accurate spatial reasoning from a single image. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high-fidelity geometry. Experiments on synthetic and real-world datasets show that PixARMesh achieves state-of-the-art reconstruction quality while producing lightweight, high-quality meshes ready for downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
Zhang, Xiang
Yoo, Sohyun
Wu, Hongrui
Li, Chuan
Xie, Jianwen
Tu, Zhuowen
Computer Vision and Pattern Recognition
Graphics
Machine Learning
We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist-ready meshes in a single forward pass. Building on recent advances in mesh generative models, we augment a point-cloud encoder with pixel-aligned image features and global scene context via cross-attention, enabling accurate spatial reasoning from a single image. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high-fidelity geometry. Experiments on synthetic and real-world datasets show that PixARMesh achieves state-of-the-art reconstruction quality while producing lightweight, high-quality meshes ready for downstream applications.
title PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2603.05888