FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D Reconstruction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tan, Guan Yuan, Vu, Ngoc Tuan, Pal, Arghya, Rajanala, Sailaja, -W., Raphael Phan C., Srinivas, Mettu, Ting, Chee-Ming
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914315645222912
author Tan, Guan Yuan
Vu, Ngoc Tuan
Pal, Arghya
Rajanala, Sailaja
-W., Raphael Phan C.
Srinivas, Mettu
Ting, Chee-Ming
author_facet Tan, Guan Yuan
Vu, Ngoc Tuan
Pal, Arghya
Rajanala, Sailaja
-W., Raphael Phan C.
Srinivas, Mettu
Ting, Chee-Ming
contents We introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron (MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D, overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations and a Global Motion Network (GMN) for capturing long-range dynamics, refined through mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08558
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D Reconstruction
Tan, Guan Yuan
Vu, Ngoc Tuan
Pal, Arghya
Rajanala, Sailaja
-W., Raphael Phan C.
Srinivas, Mettu
Ting, Chee-Ming
Computer Vision and Pattern Recognition
Computer Science and Game Theory
We introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron (MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D, overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations and a Global Motion Network (GMN) for capturing long-range dynamics, refined through mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods.
title FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D Reconstruction
topic Computer Vision and Pattern Recognition
Computer Science and Game Theory
url https://arxiv.org/abs/2602.08558