MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhowmik, Aritra, Korzhenkov, Denis, Snoek, Cees G. M., Habibian, Amirhossein, Ghafoorian, Mohsen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912663663017984
author Bhowmik, Aritra
Korzhenkov, Denis
Snoek, Cees G. M.
Habibian, Amirhossein
Ghafoorian, Mohsen
author_facet Bhowmik, Aritra
Korzhenkov, Denis
Snoek, Cees G. M.
Habibian, Amirhossein
Ghafoorian, Mohsen
contents Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural videos often entail. Recent works tackle this problem by aligning diffusion model features with those from pretrained video encoders. However, these encoders mix video appearance and dynamics into entangled features, limiting the benefit of such alignment. In this paper, we propose a motion-centric alignment framework that learns a disentangled motion subspace from a pretrained video encoder. This subspace is optimized to predict ground-truth optical flow, ensuring it captures true motion dynamics. We then align the latent features of a text-to-video diffusion model to this new subspace, enabling the generative model to internalize motion knowledge and generate more plausible videos. Our method improves the physical commonsense in a state-of-the-art video diffusion model, while preserving adherence to textual prompts, as evidenced by empirical evaluations on VideoPhy, VideoPhy2, VBench, and VBench-2.0, along with a user study.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
Bhowmik, Aritra
Korzhenkov, Denis
Snoek, Cees G. M.
Habibian, Amirhossein
Ghafoorian, Mohsen
Computer Vision and Pattern Recognition
Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural videos often entail. Recent works tackle this problem by aligning diffusion model features with those from pretrained video encoders. However, these encoders mix video appearance and dynamics into entangled features, limiting the benefit of such alignment. In this paper, we propose a motion-centric alignment framework that learns a disentangled motion subspace from a pretrained video encoder. This subspace is optimized to predict ground-truth optical flow, ensuring it captures true motion dynamics. We then align the latent features of a text-to-video diffusion model to this new subspace, enabling the generative model to internalize motion knowledge and generate more plausible videos. Our method improves the physical commonsense in a state-of-the-art video diffusion model, while preserving adherence to textual prompts, as evidenced by empirical evaluations on VideoPhy, VideoPhy2, VBench, and VBench-2.0, along with a user study.
title MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.19022