MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lei, Jiahui, Genova, Kyle, Kopanas, George, Snavely, Noah, Guibas, Leonidas
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909841147035648
author Lei, Jiahui
Genova, Kyle
Kopanas, George
Snavely, Noah
Guibas, Leonidas
author_facet Lei, Jiahui
Genova, Kyle
Kopanas, George
Snavely, Noah
Guibas, Leonidas
contents This paper addresses the challenge of learning semantically and functionally meaningful 3D motion priors from real-world videos, in order to enable prediction of future 3D scene motion from a single input image. We propose a novel pixel-aligned Motion Map (MoMap) representation for 3D scene motion, which can be generated from existing generative image models to facilitate efficient and effective motion prediction. To learn meaningful distributions over motion, we create a large-scale database of MoMaps from over 50,000 real videos and train a diffusion model on these representations. Our motion generation not only synthesizes trajectories in 3D but also suggests a new pipeline for 2D video synthesis: first generate a MoMap, then warp an image accordingly and complete the warped point-based renderings. Experimental results demonstrate that our approach generates plausible and semantically consistent 3D scene motion.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
Lei, Jiahui
Genova, Kyle
Kopanas, George
Snavely, Noah
Guibas, Leonidas
Computer Vision and Pattern Recognition
This paper addresses the challenge of learning semantically and functionally meaningful 3D motion priors from real-world videos, in order to enable prediction of future 3D scene motion from a single input image. We propose a novel pixel-aligned Motion Map (MoMap) representation for 3D scene motion, which can be generated from existing generative image models to facilitate efficient and effective motion prediction. To learn meaningful distributions over motion, we create a large-scale database of MoMaps from over 50,000 real videos and train a diffusion model on these representations. Our motion generation not only synthesizes trajectories in 3D but also suggests a new pipeline for 2D video synthesis: first generate a MoMap, then warp an image accordingly and complete the warped point-based renderings. Experimental results demonstrate that our approach generates plausible and semantically consistent 3D scene motion.
title MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.11107