Improving the Physics of Video Generation with VJEPA-2 Reward Signal

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Jianhao, Zhang, Xiaofeng, Friedrich, Felix, Beltran-Velez, Nicolas, Hall, Melissa, Askari-Hemmat, Reyhane, Han, Xiaochuang, Ballas, Nicolas, Drozdzal, Michal, Romero-Soriano, Adriana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915576140529664
author Yuan, Jianhao
Zhang, Xiaofeng
Friedrich, Felix
Beltran-Velez, Nicolas
Hall, Melissa
Askari-Hemmat, Reyhane
Han, Xiaochuang
Ballas, Nicolas
Drozdzal, Michal
Romero-Soriano, Adriana
author_facet Yuan, Jianhao
Zhang, Xiaofeng
Friedrich, Felix
Beltran-Velez, Nicolas
Hall, Melissa
Askari-Hemmat, Reyhane
Han, Xiaochuang
Ballas, Nicolas
Drozdzal, Michal
Romero-Soriano, Adriana
contents This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative models exhibit severely limited physical understanding, and often produce implausible videos. The Physics IQ benchmark has shown that visual realism does not imply physics understanding. Yet, intuitive physics understanding has shown to emerge from SSL pretraining on natural videos. In this report, we investigate whether we can leverage SSL-based video world models to improve the physics plausibility of video generative models. In particular, we build ontop of the state-of-the-art video generative model MAGI-1 and couple it with the recently introduced Video Joint Embedding Predictive Architecture 2 (VJEPA-2) to guide the generation process. We show that by leveraging VJEPA-2 as reward signal, we can improve the physics plausibility of state-of-the-art video generative models by ~6%.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving the Physics of Video Generation with VJEPA-2 Reward Signal
Yuan, Jianhao
Zhang, Xiaofeng
Friedrich, Felix
Beltran-Velez, Nicolas
Hall, Melissa
Askari-Hemmat, Reyhane
Han, Xiaochuang
Ballas, Nicolas
Drozdzal, Michal
Romero-Soriano, Adriana
Computer Vision and Pattern Recognition
Graphics
This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative models exhibit severely limited physical understanding, and often produce implausible videos. The Physics IQ benchmark has shown that visual realism does not imply physics understanding. Yet, intuitive physics understanding has shown to emerge from SSL pretraining on natural videos. In this report, we investigate whether we can leverage SSL-based video world models to improve the physics plausibility of video generative models. In particular, we build ontop of the state-of-the-art video generative model MAGI-1 and couple it with the recently introduced Video Joint Embedding Predictive Architecture 2 (VJEPA-2) to guide the generation process. We show that by leveraging VJEPA-2 as reward signal, we can improve the physics plausibility of state-of-the-art video generative models by ~6%.
title Improving the Physics of Video Generation with VJEPA-2 Reward Signal
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2510.21840