AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Yu, Xia, Menghan, Liu, Gongye, Bai, Jianhong, Wang, Xintao, Zhang, Conglang, Lin, Yuxuan, Chu, Ruihang, Wan, Pengfei, Yang, Yujiu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914089590063104
author Li, Yu
Xia, Menghan
Liu, Gongye
Bai, Jianhong
Wang, Xintao
Zhang, Conglang
Lin, Yuxuan
Chu, Ruihang
Wan, Pengfei
Yang, Yujiu
author_facet Li, Yu
Xia, Menghan
Liu, Gongye
Bai, Jianhong
Wang, Xintao
Zhang, Conglang
Lin, Yuxuan
Chu, Ruihang
Wan, Pengfei
Yang, Yujiu
contents Recent Text-to-Video (T2V) models have demonstrated powerful capability in visual simulation of real-world geometry and physical laws, indicating its potential as implicit world models. Inspired by this, we explore the feasibility of leveraging the video generation prior for viewpoint planning from given 4D scenes, since videos internally accompany dynamic scenes with natural viewpoints. To this end, we propose a two-stage paradigm to adapt pre-trained T2V models for viewpoint prediction, in a compatible manner. First, we inject the 4D scene representation into the pre-trained T2V model via an adaptive learning branch, where the 4D scene is viewpoint-agnostic and the conditional generated video embeds the viewpoints visually. Then, we formulate viewpoint extraction as a hybrid-condition guided camera extrinsic denoising process. Specifically, a camera extrinsic diffusion branch is further introduced onto the pre-trained T2V model, by taking the generated video and 4D scene as input. Experimental results show the superiority of our proposed method over existing competitors, and ablation studies validate the effectiveness of our key technical designs. To some extent, this work proves the potential of video generation models toward 4D interaction in real world.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
Li, Yu
Xia, Menghan
Liu, Gongye
Bai, Jianhong
Wang, Xintao
Zhang, Conglang
Lin, Yuxuan
Chu, Ruihang
Wan, Pengfei
Yang, Yujiu
Computer Vision and Pattern Recognition
Recent Text-to-Video (T2V) models have demonstrated powerful capability in visual simulation of real-world geometry and physical laws, indicating its potential as implicit world models. Inspired by this, we explore the feasibility of leveraging the video generation prior for viewpoint planning from given 4D scenes, since videos internally accompany dynamic scenes with natural viewpoints. To this end, we propose a two-stage paradigm to adapt pre-trained T2V models for viewpoint prediction, in a compatible manner. First, we inject the 4D scene representation into the pre-trained T2V model via an adaptive learning branch, where the 4D scene is viewpoint-agnostic and the conditional generated video embeds the viewpoints visually. Then, we formulate viewpoint extraction as a hybrid-condition guided camera extrinsic denoising process. Specifically, a camera extrinsic diffusion branch is further introduced onto the pre-trained T2V model, by taking the generated video and 4D scene as input. Experimental results show the superiority of our proposed method over existing competitors, and ablation studies validate the effectiveness of our key technical designs. To some extent, this work proves the potential of video generation models toward 4D interaction in real world.
title AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10670