GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ren, Xuanchi, Shen, Tianchang, Huang, Jiahui, Ling, Huan, Lu, Yifan, Nimier-David, Merlin, Müller, Thomas, Keller, Alexander, Fidler, Sanja, Gao, Jun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910859783045120
author Ren, Xuanchi
Shen, Tianchang
Huang, Jiahui
Ling, Huan
Lu, Yifan
Nimier-David, Merlin
Müller, Thomas
Keller, Alexander
Fidler, Sanja
Gao, Jun
author_facet Ren, Xuanchi
Shen, Tianchang
Huang, Jiahui
Ling, Huan
Lu, Yifan
Nimier-David, Merlin
Müller, Thomas
Keller, Alexander
Fidler, Sanja
Gao, Jun
contents We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if implemented at all, is imprecise, because camera parameters are mere inputs to the neural network which must then infer how the video depends on the camera. In contrast, GEN3C is guided by a 3D cache: point clouds obtained by predicting the pixel-wise depth of seed images or previously generated frames. When generating the next frames, GEN3C is conditioned on the 2D renderings of the 3D cache with the new camera trajectory provided by the user. Crucially, this means that GEN3C neither has to remember what it previously generated nor does it have to infer the image structure from the camera pose. The model, instead, can focus all its generative power on previously unobserved regions, as well as advancing the scene state to the next frame. Our results demonstrate more precise camera control than prior work, as well as state-of-the-art results in sparse-view novel view synthesis, even in challenging settings such as driving scenes and monocular dynamic video. Results are best viewed in videos. Check out our webpage! https://research.nvidia.com/labs/toronto-ai/GEN3C/
format Preprint
id arxiv_https___arxiv_org_abs_2503_03751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
Ren, Xuanchi
Shen, Tianchang
Huang, Jiahui
Ling, Huan
Lu, Yifan
Nimier-David, Merlin
Müller, Thomas
Keller, Alexander
Fidler, Sanja
Gao, Jun
Computer Vision and Pattern Recognition
Graphics
We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if implemented at all, is imprecise, because camera parameters are mere inputs to the neural network which must then infer how the video depends on the camera. In contrast, GEN3C is guided by a 3D cache: point clouds obtained by predicting the pixel-wise depth of seed images or previously generated frames. When generating the next frames, GEN3C is conditioned on the 2D renderings of the 3D cache with the new camera trajectory provided by the user. Crucially, this means that GEN3C neither has to remember what it previously generated nor does it have to infer the image structure from the camera pose. The model, instead, can focus all its generative power on previously unobserved regions, as well as advancing the scene state to the next frame. Our results demonstrate more precise camera control than prior work, as well as state-of-the-art results in sparse-view novel view synthesis, even in challenging settings such as driving scenes and monocular dynamic video. Results are best viewed in videos. Check out our webpage! https://research.nvidia.com/labs/toronto-ai/GEN3C/
title GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2503.03751