Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wimbauer, Felix, Manhardt, Fabian, Oechsle, Michael, Kalischek, Nikolai, Rupprecht, Christian, Cremers, Daniel, Tombari, Federico
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915972719312896
author Wimbauer, Felix
Manhardt, Fabian
Oechsle, Michael
Kalischek, Nikolai
Rupprecht, Christian
Cremers, Daniel
Tombari, Federico
author_facet Wimbauer, Felix
Manhardt, Fabian
Oechsle, Michael
Kalischek, Nikolai
Rupprecht, Christian
Cremers, Daniel
Tombari, Federico
contents The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and world modeling. While panoramic images have proven effective for scene initialization, existing approaches suffer from a trade-off between visual fidelity and explorability: autoregressive expansion suffers from context drift, while panoramic video generation is limited to low resolution. We present Stepper, a unified framework for text-driven immersive 3D scene synthesis that circumvents these limitations via stepwise panoramic scene expansion. Stepper leverages a novel multi-view 360° diffusion model that enables consistent, high-resolution expansion, coupled with a geometry reconstruction pipeline that enforces geometric coherence. Trained on a new large-scale, multi-view panorama dataset, Stepper achieves state-of-the-art fidelity and structural consistency, outperforming prior approaches, thereby setting a new standard for immersive scene generation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28980
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
Wimbauer, Felix
Manhardt, Fabian
Oechsle, Michael
Kalischek, Nikolai
Rupprecht, Christian
Cremers, Daniel
Tombari, Federico
Computer Vision and Pattern Recognition
The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and world modeling. While panoramic images have proven effective for scene initialization, existing approaches suffer from a trade-off between visual fidelity and explorability: autoregressive expansion suffers from context drift, while panoramic video generation is limited to low resolution. We present Stepper, a unified framework for text-driven immersive 3D scene synthesis that circumvents these limitations via stepwise panoramic scene expansion. Stepper leverages a novel multi-view 360° diffusion model that enables consistent, high-resolution expansion, coupled with a geometry reconstruction pipeline that enforces geometric coherence. Trained on a new large-scale, multi-view panorama dataset, Stepper achieves state-of-the-art fidelity and structural consistency, outperforming prior approaches, thereby setting a new standard for immersive scene generation.
title Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.28980