Video Perception Models for 3D Scene Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Rui, Zhai, Guangyao, Bauer, Zuria, Pollefeys, Marc, Tombari, Federico, Guibas, Leonidas, Huang, Gao, Engelmann, Francis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913912400642048
author Huang, Rui
Zhai, Guangyao
Bauer, Zuria
Pollefeys, Marc
Tombari, Federico
Guibas, Leonidas
Huang, Gao
Engelmann, Francis
author_facet Huang, Rui
Zhai, Guangyao
Bauer, Zuria
Pollefeys, Marc
Tombari, Federico
Guibas, Leonidas
Huang, Gao
Engelmann, Francis
contents Traditionally, 3D scene synthesis requires expert knowledge and significant manual effort. Automating this process could greatly benefit fields such as architectural design, robotics simulation, virtual reality, and gaming. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or strong visual priors of modern image generation models. However, current LLMs demonstrate limited 3D spatial reasoning ability, which restricts their ability to generate realistic and coherent 3D scenes. Meanwhile, image generation-based methods often suffer from constraints in viewpoint selection and multi-view inconsistencies. In this work, we present Video Perception models for 3D Scene synthesis (VIPScene), a novel framework that exploits the encoded commonsense knowledge of the 3D physical world in video generation models to ensure coherent scene layouts and consistent object placements across views. VIPScene accepts both text and image prompts and seamlessly integrates video generation, feedforward 3D reconstruction, and open-vocabulary perception models to semantically and geometrically analyze each object in a scene. This enables flexible scene synthesis with high realism and structural consistency. For more precise analysis, we further introduce First-Person View Score (FPVScore) for coherence and plausibility evaluation, utilizing continuous first-person perspective to capitalize on the reasoning ability of multimodal large language models. Extensive experiments show that VIPScene significantly outperforms existing methods and generalizes well across diverse scenarios. The code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20601
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Perception Models for 3D Scene Synthesis
Huang, Rui
Zhai, Guangyao
Bauer, Zuria
Pollefeys, Marc
Tombari, Federico
Guibas, Leonidas
Huang, Gao
Engelmann, Francis
Computer Vision and Pattern Recognition
Traditionally, 3D scene synthesis requires expert knowledge and significant manual effort. Automating this process could greatly benefit fields such as architectural design, robotics simulation, virtual reality, and gaming. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or strong visual priors of modern image generation models. However, current LLMs demonstrate limited 3D spatial reasoning ability, which restricts their ability to generate realistic and coherent 3D scenes. Meanwhile, image generation-based methods often suffer from constraints in viewpoint selection and multi-view inconsistencies. In this work, we present Video Perception models for 3D Scene synthesis (VIPScene), a novel framework that exploits the encoded commonsense knowledge of the 3D physical world in video generation models to ensure coherent scene layouts and consistent object placements across views. VIPScene accepts both text and image prompts and seamlessly integrates video generation, feedforward 3D reconstruction, and open-vocabulary perception models to semantically and geometrically analyze each object in a scene. This enables flexible scene synthesis with high realism and structural consistency. For more precise analysis, we further introduce First-Person View Score (FPVScore) for coherence and plausibility evaluation, utilizing continuous first-person perspective to capitalize on the reasoning ability of multimodal large language models. Extensive experiments show that VIPScene significantly outperforms existing methods and generalizes well across diverse scenarios. The code will be released.
title Video Perception Models for 3D Scene Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.20601