LiveStre4m: Feed-Forward Live Streaming of Novel Views from Unposed Multi-View Video

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Quesado, Pedro, Akdag, Erkut, Kashefbahrami, Yasaman, Menu, Willem, Bondarev, Egor
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910112129482752
author Quesado, Pedro
Akdag, Erkut
Kashefbahrami, Yasaman
Menu, Willem
Bondarev, Egor
author_facet Quesado, Pedro
Akdag, Erkut
Kashefbahrami, Yasaman
Menu, Willem
Bondarev, Egor
contents Live-streaming Novel View Synthesis (NVS) from unposed multi-view video remains an open challenge in a wide range of applications. Existing methods for dynamic scene representation typically require ground-truth camera parameters and involve lengthy optimizations ($\approx 2.67$s), which makes them unsuitable for live streaming scenarios. To address this issue, we propose a novel viewpoint video live-streaming method (LiveStre4m), a feed-forward model for real-time NVS from unposed sparse multi-view inputs. LiveStre4m introduces a multi-view vision transformer for keyframe 3D scene reconstruction coupled with a diffusion-transformer interpolation module that ensures temporal consistency and stable streaming. In addition, a Camera Pose Predictor module is proposed to efficiently estimate both poses and intrinsics directly from RGB images, removing the reliance on known camera calibration information. Our approach enables temporally consistent novel-view video streaming in real-time using as few as two synchronized unposed input streams. LiveStre4m attains an average reconstruction time of $ 0.07$s per-frame at $ 1024 \times 768$ resolution, outperforming the optimization-based dynamic scene representation methods by orders of magnitude in runtime. These results demonstrate that LiveStre4m makes real-time NVS streaming feasible in practical settings, marking a substantial step toward deployable live novel-view synthesis systems. Code available at: https://github.com/pedro-quesado/LiveStre4m
format Preprint
id arxiv_https___arxiv_org_abs_2604_06740
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LiveStre4m: Feed-Forward Live Streaming of Novel Views from Unposed Multi-View Video
Quesado, Pedro
Akdag, Erkut
Kashefbahrami, Yasaman
Menu, Willem
Bondarev, Egor
Computer Vision and Pattern Recognition
Live-streaming Novel View Synthesis (NVS) from unposed multi-view video remains an open challenge in a wide range of applications. Existing methods for dynamic scene representation typically require ground-truth camera parameters and involve lengthy optimizations ($\approx 2.67$s), which makes them unsuitable for live streaming scenarios. To address this issue, we propose a novel viewpoint video live-streaming method (LiveStre4m), a feed-forward model for real-time NVS from unposed sparse multi-view inputs. LiveStre4m introduces a multi-view vision transformer for keyframe 3D scene reconstruction coupled with a diffusion-transformer interpolation module that ensures temporal consistency and stable streaming. In addition, a Camera Pose Predictor module is proposed to efficiently estimate both poses and intrinsics directly from RGB images, removing the reliance on known camera calibration information. Our approach enables temporally consistent novel-view video streaming in real-time using as few as two synchronized unposed input streams. LiveStre4m attains an average reconstruction time of $ 0.07$s per-frame at $ 1024 \times 768$ resolution, outperforming the optimization-based dynamic scene representation methods by orders of magnitude in runtime. These results demonstrate that LiveStre4m makes real-time NVS streaming feasible in practical settings, marking a substantial step toward deployable live novel-view synthesis systems. Code available at: https://github.com/pedro-quesado/LiveStre4m
title LiveStre4m: Feed-Forward Live Streaming of Novel Views from Unposed Multi-View Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.06740