ShapeR: Robust Conditional 3D Shape Generation from Casual Captures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siddiqui, Yawar, Frost, Duncan, Aroudj, Samir, Avetisyan, Armen, Howard-Jenkins, Henry, DeTone, Daniel, Moulon, Pierre, Wu, Qirui, Li, Zhengqin, Straub, Julian, Newcombe, Richard, Engel, Jakob
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911380575092736
author Siddiqui, Yawar
Frost, Duncan
Aroudj, Samir
Avetisyan, Armen
Howard-Jenkins, Henry
DeTone, Daniel
Moulon, Pierre
Wu, Qirui
Li, Zhengqin
Straub, Julian
Newcombe, Richard
Engel, Jakob
author_facet Siddiqui, Yawar
Frost, Duncan
Aroudj, Samir
Avetisyan, Armen
Howard-Jenkins, Henry
DeTone, Daniel
Moulon, Pierre
Wu, Qirui
Li, Zhengqin
Straub, Julian
Newcombe, Richard
Engel, Jakob
contents Recent advances in 3D shape generation have achieved impressive results, but most existing methods rely on clean, unoccluded, and well-segmented inputs. Such conditions are rarely met in real-world scenarios. We present ShapeR, a novel approach for conditional 3D object shape generation from casually captured sequences. Given an image sequence, we leverage off-the-shelf visual-inertial SLAM, 3D detection algorithms, and vision-language models to extract, for each object, a set of sparse SLAM points, posed multi-view images, and machine-generated captions. A rectified flow transformer trained to effectively condition on these modalities then generates high-fidelity metric 3D shapes. To ensure robustness to the challenges of casually captured data, we employ a range of techniques including on-the-fly compositional augmentations, a curriculum training scheme spanning object- and scene-level datasets, and strategies to handle background clutter. Additionally, we introduce a new evaluation benchmark comprising 178 in-the-wild objects across 7 real-world scenes with geometry annotations. Experiments show that ShapeR significantly outperforms existing approaches in this challenging setting, achieving an improvement of 2.7x in Chamfer distance compared to state of the art.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11514
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
Siddiqui, Yawar
Frost, Duncan
Aroudj, Samir
Avetisyan, Armen
Howard-Jenkins, Henry
DeTone, Daniel
Moulon, Pierre
Wu, Qirui
Li, Zhengqin
Straub, Julian
Newcombe, Richard
Engel, Jakob
Computer Vision and Pattern Recognition
Machine Learning
Recent advances in 3D shape generation have achieved impressive results, but most existing methods rely on clean, unoccluded, and well-segmented inputs. Such conditions are rarely met in real-world scenarios. We present ShapeR, a novel approach for conditional 3D object shape generation from casually captured sequences. Given an image sequence, we leverage off-the-shelf visual-inertial SLAM, 3D detection algorithms, and vision-language models to extract, for each object, a set of sparse SLAM points, posed multi-view images, and machine-generated captions. A rectified flow transformer trained to effectively condition on these modalities then generates high-fidelity metric 3D shapes. To ensure robustness to the challenges of casually captured data, we employ a range of techniques including on-the-fly compositional augmentations, a curriculum training scheme spanning object- and scene-level datasets, and strategies to handle background clutter. Additionally, we introduce a new evaluation benchmark comprising 178 in-the-wild objects across 7 real-world scenes with geometry annotations. Experiments show that ShapeR significantly outperforms existing approaches in this challenging setting, achieving an improvement of 2.7x in Chamfer distance compared to state of the art.
title ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2601.11514