LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Mingyang, Khan, Numair, Wang, Tianfu, Dhingra, Naina, Nam, Seonghyeon, Yang, Haitao, Hui, Zhuo, Metzler, Christopher, Vedaldi, Andrea, Pirsiavash, Hamed, Luo, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918425683558400
author Xie, Mingyang
Khan, Numair
Wang, Tianfu
Dhingra, Naina
Nam, Seonghyeon
Yang, Haitao
Hui, Zhuo
Metzler, Christopher
Vedaldi, Andrea
Pirsiavash, Hamed
Luo, Lei
author_facet Xie, Mingyang
Khan, Numair
Wang, Tianfu
Dhingra, Naina
Nam, Seonghyeon
Yang, Haitao
Hui, Zhuo
Metzler, Christopher
Vedaldi, Andrea
Pirsiavash, Hamed
Luo, Lei
contents Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to drift and deformation under viewpoint changes. On the other hand, geometrically-conditioned models depend on estimated depth and explicit reconstruction, making them susceptible to depth inaccuracies and calibration errors. We propose to address these challenges by using the implicit geometric knowledge embedded in the latent space of a large 4D reconstruction model to condition the video generation process. These latents capture scene structure in a continuous space without explicit reconstruction. Therefore, they provide a flexible representation that allows the pretrained diffusion prior to regularize errors more effectively. By jointly conditioning on these latents and source camera poses, we demonstrate that our model achieves state-of-the-art results on the video re-rendering task. Project webpage is https://lavr-4d-scene-rerender.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14674
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
Xie, Mingyang
Khan, Numair
Wang, Tianfu
Dhingra, Naina
Nam, Seonghyeon
Yang, Haitao
Hui, Zhuo
Metzler, Christopher
Vedaldi, Andrea
Pirsiavash, Hamed
Luo, Lei
Computer Vision and Pattern Recognition
Machine Learning
Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to drift and deformation under viewpoint changes. On the other hand, geometrically-conditioned models depend on estimated depth and explicit reconstruction, making them susceptible to depth inaccuracies and calibration errors. We propose to address these challenges by using the implicit geometric knowledge embedded in the latent space of a large 4D reconstruction model to condition the video generation process. These latents capture scene structure in a continuous space without explicit reconstruction. Therefore, they provide a flexible representation that allows the pretrained diffusion prior to regularize errors more effectively. By jointly conditioning on these latents and source camera poses, we demonstrate that our model achieves state-of-the-art results on the video re-rendering task. Project webpage is https://lavr-4d-scene-rerender.github.io/.
title LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2601.14674