Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Mungyeom, Jeon, Minkyeong, An, Honggyu, Jung, Jaewoo, Ko, Hyuna, Han, Jisang, Yu, Hyeonseo, Shin, Donghwan, Hong, Sunghwan, Narihira, Takuya, Fukuda, Kazumi, Mitsufuji, Yuki, Kim, Seungryong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916066834251776
author Kim, Mungyeom
Jeon, Minkyeong
An, Honggyu
Jung, Jaewoo
Ko, Hyuna
Han, Jisang
Yu, Hyeonseo
Shin, Donghwan
Hong, Sunghwan
Narihira, Takuya
Fukuda, Kazumi
Mitsufuji, Yuki
Kim, Seungryong
author_facet Kim, Mungyeom
Jeon, Minkyeong
An, Honggyu
Jung, Jaewoo
Ko, Hyuna
Han, Jisang
Yu, Hyeonseo
Shin, Donghwan
Hong, Sunghwan
Narihira, Takuya
Fukuda, Kazumi
Mitsufuji, Yuki
Kim, Seungryong
contents Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame, suffering from duplicated Gaussians and view-dependent biases that hinder effective learning of scene motion. We present C4G, a feed-forward 4D reconstruction framework built upon a compact set of timestamp-conditioned learnable Gaussian query tokens. Each token aggregates corresponding features across the full temporal context and decodes a 3D Gaussian whose position is modulated by the target timestamp, enabling globally coherent motion modeling without per-scene optimization. To capture fine-grained details, we further introduce a video diffusion model-based rendering enhancement module. Since our framework effectively aggregates features into Gaussians, we extend this capability to feature lifting, producing a 4D feature field that supports point tracking and dynamic scene understanding. C4G achieves strong novel-view synthesis performance using significantly fewer Gaussians and without requiring camera poses, while exhibiting stronger motion modeling and robustness to large temporal gaps.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31595
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
Kim, Mungyeom
Jeon, Minkyeong
An, Honggyu
Jung, Jaewoo
Ko, Hyuna
Han, Jisang
Yu, Hyeonseo
Shin, Donghwan
Hong, Sunghwan
Narihira, Takuya
Fukuda, Kazumi
Mitsufuji, Yuki
Kim, Seungryong
Computer Vision and Pattern Recognition
Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians pixel-wise for each frame, suffering from duplicated Gaussians and view-dependent biases that hinder effective learning of scene motion. We present C4G, a feed-forward 4D reconstruction framework built upon a compact set of timestamp-conditioned learnable Gaussian query tokens. Each token aggregates corresponding features across the full temporal context and decodes a 3D Gaussian whose position is modulated by the target timestamp, enabling globally coherent motion modeling without per-scene optimization. To capture fine-grained details, we further introduce a video diffusion model-based rendering enhancement module. Since our framework effectively aggregates features into Gaussians, we extend this capability to feature lifting, producing a 4D feature field that supports point tracking and dynamic scene understanding. C4G achieves strong novel-view synthesis performance using significantly fewer Gaussians and without requiring camera poses, while exhibiting stronger motion modeling and robustness to large temporal gaps.
title Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.31595