Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Shuo, Artan, Unal, Mielle, Malcolm, Lilienthaland, Achim J., Magnusson, Martin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910052383719424
author Sun, Shuo
Artan, Unal
Mielle, Malcolm
Lilienthaland, Achim J.
Magnusson, Martin
author_facet Sun, Shuo
Artan, Unal
Mielle, Malcolm
Lilienthaland, Achim J.
Magnusson, Martin
contents We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches either handle only single-camera input or require rigidly mounted, pre-calibrated camera rigs, limiting their practical applicability. We propose a two-stage optimization framework that decouples the task into robust camera tracking and dense depth refinement. In the first stage, we extend single-camera visual SLAM to the multi-camera setting by constructing a spatiotemporal connection graph that exploits both intra-camera temporal continuity and inter-camera spatial overlap, enabling consistent scale and robust tracking. To ensure robustness under limited overlap, we introduce a wide-baseline initialization strategy using feed-forward reconstruction models. In the second stage, we refine depth and camera poses by optimizing dense inter- and intra-camera consistency using wide-baseline optical flow. Additionally, we introduce MultiCamRobolab, a new real-world dataset with ground-truth poses from a motion capture system. Finally, we demonstrate that our method significantly outperforms state-of-the-art feed-forward models on both synthetic and real-world benchmarks, while requiring less memory.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12064
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
Sun, Shuo
Artan, Unal
Mielle, Malcolm
Lilienthaland, Achim J.
Magnusson, Martin
Computer Vision and Pattern Recognition
We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches either handle only single-camera input or require rigidly mounted, pre-calibrated camera rigs, limiting their practical applicability. We propose a two-stage optimization framework that decouples the task into robust camera tracking and dense depth refinement. In the first stage, we extend single-camera visual SLAM to the multi-camera setting by constructing a spatiotemporal connection graph that exploits both intra-camera temporal continuity and inter-camera spatial overlap, enabling consistent scale and robust tracking. To ensure robustness under limited overlap, we introduce a wide-baseline initialization strategy using feed-forward reconstruction models. In the second stage, we refine depth and camera poses by optimizing dense inter- and intra-camera consistency using wide-baseline optical flow. Additionally, we introduce MultiCamRobolab, a new real-world dataset with ground-truth poses from a motion capture system. Finally, we demonstrate that our method significantly outperforms state-of-the-art feed-forward models on both synthetic and real-world benchmarks, while requiring less memory.
title Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12064