Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qian, Shenhan, Zhang, Ganlin, Wu, Shangzhe, Cremers, Daniel
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912906179772416
author Qian, Shenhan
Zhang, Ganlin
Wu, Shangzhe
Cremers, Daniel
author_facet Qian, Shenhan
Zhang, Ganlin
Wu, Shangzhe
Cremers, Daniel
contents Reconstructing and tracking dynamic 3D scenes remains a fundamental challenge in computer vision. Existing approaches often decouple geometry from motion: multi-view reconstruction methods assume static scenes, while dynamic tracking frameworks rely on explicit camera pose estimation or separate motion models. We propose Flow4R, a unified framework that treats camera-space scene flow as the central representation linking 3D structure, object motion, and camera motion. Flow4R predicts a minimal per-pixel property set-3D point position, scene flow, pose weight, and confidence-from two-view inputs using a Vision Transformer. This flow-centric formulation allows local geometry and bidirectional motion to be inferred symmetrically with a shared decoder in a single forward pass, without requiring explicit pose regressors or bundle adjustment. Trained jointly on static and dynamic datasets, Flow4R achieves state-of-the-art performance on 4D reconstruction and tracking tasks, demonstrating the effectiveness of the flow-central representation for spatiotemporal scene understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14021
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
Qian, Shenhan
Zhang, Ganlin
Wu, Shangzhe
Cremers, Daniel
Computer Vision and Pattern Recognition
Reconstructing and tracking dynamic 3D scenes remains a fundamental challenge in computer vision. Existing approaches often decouple geometry from motion: multi-view reconstruction methods assume static scenes, while dynamic tracking frameworks rely on explicit camera pose estimation or separate motion models. We propose Flow4R, a unified framework that treats camera-space scene flow as the central representation linking 3D structure, object motion, and camera motion. Flow4R predicts a minimal per-pixel property set-3D point position, scene flow, pose weight, and confidence-from two-view inputs using a Vision Transformer. This flow-centric formulation allows local geometry and bidirectional motion to be inferred symmetrically with a shared decoder in a single forward pass, without requiring explicit pose regressors or bundle adjustment. Trained jointly on static and dynamic datasets, Flow4R achieves state-of-the-art performance on 4D reconstruction and tracking tasks, demonstrating the effectiveness of the flow-central representation for spatiotemporal scene understanding.
title Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.14021