GRVS: a Generalizable and Recurrent Approach to Monocular Dynamic View Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tanay, Thomas, Brahimi, Mohammed, Nazarczuk, Michal, Zhang, Qingwen, Catley-Chandar, Sibi, Moreau, Arthur, Zhang, Zhensong, Pérez-Pellitero, Eduardo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911557906071552
author Tanay, Thomas
Brahimi, Mohammed
Nazarczuk, Michal
Zhang, Qingwen
Catley-Chandar, Sibi
Moreau, Arthur
Zhang, Zhensong
Pérez-Pellitero, Eduardo
author_facet Tanay, Thomas
Brahimi, Mohammed
Nazarczuk, Michal
Zhang, Qingwen
Catley-Chandar, Sibi
Moreau, Arthur
Zhang, Zhensong
Pérez-Pellitero, Eduardo
contents Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view information is hard to exploit. Diffusion-based approaches that integrate camera control into large pre-trained models can produce visually plausible videos but frequently suffer from geometric inconsistencies across both static and dynamic areas. Both families of methods also require substantial computational resources. Building on the success of generalizable models for static novel view synthesis, we adapt the framework to dynamic inputs and propose a new model with two key components: (1) a recurrent loop that enables unbounded and asynchronous mapping between input and target videos and (2) an efficient use of plane sweeps over dynamic inputs to disentangle camera and scene motion, and achieve fine-grained, six-degrees-of-freedom camera controls. We train and evaluate our model on the UCSD dataset and on Kubric-4D-dyn, a new monocular dynamic dataset featuring longer, higher resolution sequences with more complex scene dynamics than existing alternatives. Our model outperforms four Gaussian Splatting-based scene-specific approaches, as well as two diffusion-based approaches in reconstructing fine-grained geometric details across both static and dynamic regions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29734
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GRVS: a Generalizable and Recurrent Approach to Monocular Dynamic View Synthesis
Tanay, Thomas
Brahimi, Mohammed
Nazarczuk, Michal
Zhang, Qingwen
Catley-Chandar, Sibi
Moreau, Arthur
Zhang, Zhensong
Pérez-Pellitero, Eduardo
Computer Vision and Pattern Recognition
Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view information is hard to exploit. Diffusion-based approaches that integrate camera control into large pre-trained models can produce visually plausible videos but frequently suffer from geometric inconsistencies across both static and dynamic areas. Both families of methods also require substantial computational resources. Building on the success of generalizable models for static novel view synthesis, we adapt the framework to dynamic inputs and propose a new model with two key components: (1) a recurrent loop that enables unbounded and asynchronous mapping between input and target videos and (2) an efficient use of plane sweeps over dynamic inputs to disentangle camera and scene motion, and achieve fine-grained, six-degrees-of-freedom camera controls. We train and evaluate our model on the UCSD dataset and on Kubric-4D-dyn, a new monocular dynamic dataset featuring longer, higher resolution sequences with more complex scene dynamics than existing alternatives. Our model outperforms four Gaussian Splatting-based scene-specific approaches, as well as two diffusion-based approaches in reconstructing fine-grained geometric details across both static and dynamic regions.
title GRVS: a Generalizable and Recurrent Approach to Monocular Dynamic View Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.29734