StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Yunzhi, Xu, Zhen, Lin, Haotong, Jin, Haian, Guo, Haoyu, Wang, Yida, Zhan, Kun, Lang, Xianpeng, Bao, Hujun, Zhou, Xiaowei, Peng, Sida
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918130379390976
author Yan, Yunzhi
Xu, Zhen
Lin, Haotong
Jin, Haian
Guo, Haoyu
Wang, Yida
Zhan, Kun
Lang, Xianpeng
Bao, Hujun
Zhou, Xiaowei
Peng, Sida
author_facet Yan, Yunzhi
Xu, Zhen
Lin, Haotong
Jin, Haian
Guo, Haoyu
Wang, Yida
Zhan, Kun
Lang, Xianpeng
Bao, Hujun
Zhou, Xiaowei
Peng, Sida
contents This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the performance significantly degrades as the viewpoint deviates from the training trajectory. To mitigate this problem, we introduce StreetCrafter, a novel controllable video diffusion model that utilizes LiDAR point cloud renderings as pixel-level conditions, which fully exploits the generative prior for novel view synthesis, while preserving precise camera control. Moreover, the utilization of pixel-level LiDAR conditions allows us to make accurate pixel-level edits to target scenes. In addition, the generative prior of StreetCrafter can be effectively incorporated into dynamic scene representations to achieve real-time rendering. Experiments on Waymo Open Dataset and PandaSet demonstrate that our model enables flexible control over viewpoint changes, enlarging the view synthesis regions for satisfying rendering, which outperforms existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13188
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
Yan, Yunzhi
Xu, Zhen
Lin, Haotong
Jin, Haian
Guo, Haoyu
Wang, Yida
Zhan, Kun
Lang, Xianpeng
Bao, Hujun
Zhou, Xiaowei
Peng, Sida
Computer Vision and Pattern Recognition
This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the performance significantly degrades as the viewpoint deviates from the training trajectory. To mitigate this problem, we introduce StreetCrafter, a novel controllable video diffusion model that utilizes LiDAR point cloud renderings as pixel-level conditions, which fully exploits the generative prior for novel view synthesis, while preserving precise camera control. Moreover, the utilization of pixel-level LiDAR conditions allows us to make accurate pixel-level edits to target scenes. In addition, the generative prior of StreetCrafter can be effectively incorporated into dynamic scene representations to achieve real-time rendering. Experiments on Waymo Open Dataset and PandaSet demonstrate that our model enables flexible control over viewpoint changes, enlarging the view synthesis regions for satisfying rendering, which outperforms existing methods.
title StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13188