VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918375319404544 |
|---|---|
| author | Zhu, Muhua Jin, Xinhao Zhang, Yu Xue, Yifei Ji, Tie Lao, Yizhen |
| author_facet | Zhu, Muhua Jin, Xinhao Zhang, Yu Xue, Yifei Ji, Tie Lao, Yizhen |
| contents | Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile optimization pipelines that fail under extreme motions. To bridge this gap, we propose VS3R, a framework that synergizes feed-forward 3D reconstruction with generative video diffusion. Our pipeline jointly estimates camera parameters, depth, and masks to ensure all-scenario reliability, and introduces a Hybrid Stabilized Rendering module that fuses semantic and geometric cues for dynamic consistency. Finally, a Dual-Stream Video Diffusion Model restores disoccluded regions and rectifies artifacts by synergizing structural guidance with semantic anchors. Collectively, VS3R achieves high-fidelity, full-frame stabilization across diverse camera models and significantly outperforms state-of-the-art methods in robustness and visual quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_05851 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction Zhu, Muhua Jin, Xinhao Zhang, Yu Xue, Yifei Ji, Tie Lao, Yizhen Computer Vision and Pattern Recognition Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2D methods suffer from aggressive cropping, 3D techniques are often undermined by fragile optimization pipelines that fail under extreme motions. To bridge this gap, we propose VS3R, a framework that synergizes feed-forward 3D reconstruction with generative video diffusion. Our pipeline jointly estimates camera parameters, depth, and masks to ensure all-scenario reliability, and introduces a Hybrid Stabilized Rendering module that fuses semantic and geometric cues for dynamic consistency. Finally, a Dual-Stream Video Diffusion Model restores disoccluded regions and rectifies artifacts by synergizing structural guidance with semantic anchors. Collectively, VS3R achieves high-fidelity, full-frame stabilization across diverse camera models and significantly outperforms state-of-the-art methods in robustness and visual quality. |
| title | VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.05851 |