NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yu, Wang, Xijun, Wickremasinghe, Tharindu, Nadir, Zeeshan, Ma, Bole, Chan, Stanley H.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914428280111104
author Yuan, Yu
Wang, Xijun
Wickremasinghe, Tharindu
Nadir, Zeeshan
Ma, Bole
Chan, Stanley H.
author_facet Yuan, Yu
Wang, Xijun
Wickremasinghe, Tharindu
Nadir, Zeeshan
Ma, Bole
Chan, Stanley H.
contents A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack precise parameter control, struggling to generate physically consistent dynamics under different initial conditions. We argue that this fundamental limitation stems from current models learning motion distributions solely from appearance, while lacking an understanding of the underlying dynamics. In this work, we propose NewtonGen, a framework that integrates data-driven synthesis with learnable physical principles. At its core lies trainable Neural Newtonian Dynamics (NND), which can model and predict a variety of Newtonian motions, thereby injecting latent dynamical constraints into the video generation process. By jointly leveraging data priors and dynamical guidance, NewtonGen enables physically consistent video synthesis with precise parameter control. All data and code are available at https://github.com/pandayuanyu/NewtonGen
format Preprint
id arxiv_https___arxiv_org_abs_2509_21309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
Yuan, Yu
Wang, Xijun
Wickremasinghe, Tharindu
Nadir, Zeeshan
Ma, Bole
Chan, Stanley H.
Computer Vision and Pattern Recognition
A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack precise parameter control, struggling to generate physically consistent dynamics under different initial conditions. We argue that this fundamental limitation stems from current models learning motion distributions solely from appearance, while lacking an understanding of the underlying dynamics. In this work, we propose NewtonGen, a framework that integrates data-driven synthesis with learnable physical principles. At its core lies trainable Neural Newtonian Dynamics (NND), which can model and predict a variety of Newtonian motions, thereby injecting latent dynamical constraints into the video generation process. By jointly leveraging data priors and dynamical guidance, NewtonGen enables physically consistent video synthesis with precise parameter control. All data and code are available at https://github.com/pandayuanyu/NewtonGen
title NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.21309