VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chefer, Hila, Singer, Uriel, Zohar, Amit, Kirstain, Yuval, Polyak, Adam, Taigman, Yaniv, Wolf, Lior, Sheynin, Shelly
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908379784413184
author Chefer, Hila
Singer, Uriel
Zohar, Amit
Kirstain, Yuval
Polyak, Adam
Taigman, Yaniv
Wolf, Lior
Sheynin, Shelly
author_facet Chefer, Hila
Singer, Uriel
Zohar, Amit
Kirstain, Yuval
Polyak, Adam
Taigman, Yaniv
Wolf, Lior
Sheynin, Shelly
contents Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2502_02492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
Chefer, Hila
Singer, Uriel
Zohar, Amit
Kirstain, Yuval
Polyak, Adam
Taigman, Yaniv
Wolf, Lior
Sheynin, Shelly
Computer Vision and Pattern Recognition
Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/
title VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.02492