PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918375405387776 |
|---|---|
| author | Zhang, Xiang Yoo, Sohyun Wu, Hongrui Li, Chuan Xie, Jianwen Tu, Zhuowen |
| author_facet | Zhang, Xiang Yoo, Sohyun Wu, Hongrui Li, Chuan Xie, Jianwen Tu, Zhuowen |
| contents | We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist-ready meshes in a single forward pass. Building on recent advances in mesh generative models, we augment a point-cloud encoder with pixel-aligned image features and global scene context via cross-attention, enabling accurate spatial reasoning from a single image. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high-fidelity geometry. Experiments on synthetic and real-world datasets show that PixARMesh achieves state-of-the-art reconstruction quality while producing lightweight, high-quality meshes ready for downstream applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_05888 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction Zhang, Xiang Yoo, Sohyun Wu, Hongrui Li, Chuan Xie, Jianwen Tu, Zhuowen Computer Vision and Pattern Recognition Graphics Machine Learning We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed distance fields and post-hoc layout optimization, PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist-ready meshes in a single forward pass. Building on recent advances in mesh generative models, we augment a point-cloud encoder with pixel-aligned image features and global scene context via cross-attention, enabling accurate spatial reasoning from a single image. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high-fidelity geometry. Experiments on synthetic and real-world datasets show that PixARMesh achieves state-of-the-art reconstruction quality while producing lightweight, high-quality meshes ready for downstream applications. |
| title | PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction |
| topic | Computer Vision and Pattern Recognition Graphics Machine Learning |
| url | https://arxiv.org/abs/2603.05888 |