VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Muhua, Jin, Xinhao, Zhang, Yu, Xue, Yifei, Ji, Tie, Lao, Yizhen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918375319404544
author Zhu, Muhua
Jin, Xinhao
Zhang, Yu
Xue, Yifei
Ji, Tie
Lao, Yizhen
author_facet Zhu, Muhua
Jin, Xinhao
Zhang, Yu
Xue, Yifei
Ji, Tie
Lao, Yizhen
contents Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile optimization pipelines that fail under extreme motions. To bridge this gap, we propose VS3R, a framework that synergizes feed-forward 3D reconstruction with generative video diffusion. Our pipeline jointly estimates camera parameters, depth, and masks to ensure all-scenario reliability, and introduces a Hybrid Stabilized Rendering module that fuses semantic and geometric cues for dynamic consistency. Finally, a Dual-Stream Video Diffusion Model restores disoccluded regions and rectifies artifacts by synergizing structural guidance with semantic anchors. Collectively, VS3R achieves high-fidelity, full-frame stabilization across diverse camera models and significantly outperforms state-of-the-art methods in robustness and visual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05851
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction
Zhu, Muhua
Jin, Xinhao
Zhang, Yu
Xue, Yifei
Ji, Tie
Lao, Yizhen
Computer Vision and Pattern Recognition
Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile optimization pipelines that fail under extreme motions. To bridge this gap, we propose VS3R, a framework that synergizes feed-forward 3D reconstruction with generative video diffusion. Our pipeline jointly estimates camera parameters, depth, and masks to ensure all-scenario reliability, and introduces a Hybrid Stabilized Rendering module that fuses semantic and geometric cues for dynamic consistency. Finally, a Dual-Stream Video Diffusion Model restores disoccluded regions and rectifies artifacts by synergizing structural guidance with semantic anchors. Collectively, VS3R achieves high-fidelity, full-frame stabilization across diverse camera models and significantly outperforms state-of-the-art methods in robustness and visual quality.
title VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.05851