FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Luxi, Zhou, Zihan, Zhao, Min, Wang, Yikai, Zhang, Ge, Huang, Wenhao, Sun, Hao, Wen, Ji-Rong, Li, Chongxuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916656527179776
author Chen, Luxi
Zhou, Zihan
Zhao, Min
Wang, Yikai
Zhang, Ge
Huang, Wenhao
Sun, Hao
Wen, Ji-Rong
Li, Chongxuan
author_facet Chen, Luxi
Zhou, Zihan
Zhao, Min
Wang, Yikai
Zhang, Ge
Huang, Wenhao
Sun, Hao
Wen, Ji-Rong
Li, Chongxuan
contents Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong video-to-video (V2V) diffusion model to generate high-quality novel view images from incomplete input rendered from a coarse scene, and (2) a progressive expansion process to construct a complete 3D scene. In particular, leveraging an advanced pre-trained video model and accurate depth-estimated training pairs, our V2V model can generate novel views under large camera pose variations. Building upon it, FlexWorld progressively generates new 3D content and integrates it into the global scene through geometry-aware scene fusion. Extensive experiments demonstrate the effectiveness of FlexWorld in generating high-quality novel view videos and flexible-view 3D scenes from single images, achieving superior visual quality under multiple popular metrics and datasets compared to existing state-of-the-art methods. Qualitatively, we highlight that FlexWorld can generate high-fidelity scenes with flexible views like 360° rotations and zooming. Project page: https://ml-gsai.github.io/FlexWorld.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13265
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis
Chen, Luxi
Zhou, Zihan
Zhao, Min
Wang, Yikai
Zhang, Ge
Huang, Wenhao
Sun, Hao
Wen, Ji-Rong
Li, Chongxuan
Computer Vision and Pattern Recognition
Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong video-to-video (V2V) diffusion model to generate high-quality novel view images from incomplete input rendered from a coarse scene, and (2) a progressive expansion process to construct a complete 3D scene. In particular, leveraging an advanced pre-trained video model and accurate depth-estimated training pairs, our V2V model can generate novel views under large camera pose variations. Building upon it, FlexWorld progressively generates new 3D content and integrates it into the global scene through geometry-aware scene fusion. Extensive experiments demonstrate the effectiveness of FlexWorld in generating high-quality novel view videos and flexible-view 3D scenes from single images, achieving superior visual quality under multiple popular metrics and datasets compared to existing state-of-the-art methods. Qualitatively, we highlight that FlexWorld can generate high-fidelity scenes with flexible views like 360° rotations and zooming. Project page: https://ml-gsai.github.io/FlexWorld.
title FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.13265