Moaw: Unleashing Motion Awareness for Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tianqi, Wang, Ziyi, Zheng, Wenzhao, Chen, Weiliang, Huang, Yuanhui, Huang, Zhengyang, Zhou, Jie, Lu, Jiwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917209698205696
author Zhang, Tianqi
Wang, Ziyi
Zheng, Wenzhao
Chen, Weiliang
Huang, Yuanhui
Huang, Zhengyang
Zhou, Jie
Lu, Jiwen
author_facet Zhang, Tianqi
Wang, Ziyi
Zheng, Wenzhao
Chen, Weiliang
Huang, Yuanhui
Huang, Zhengyang
Zhou, Jie
Lu, Jiwen
contents Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot setting. Motivated by these findings, we investigate whether supervised training can more fully harness the tracking capability of video diffusion models. To this end, we propose Moaw, a framework that unleashes motion awareness for video diffusion models and leverages it to facilitate motion transfer. Specifically, we train a diffusion model for motion perception, shifting its modality from image-to-video generation to video-to-dense-tracking. We then construct a motion-labeled dataset to identify features that encode the strongest motion information, and inject them into a structurally identical video generation model. Owing to the homogeneity between the two networks, these features can be naturally adapted in a zero-shot manner, enabling motion transfer without additional adapters. Our work provides a new paradigm for bridging generative modeling and motion understanding, paving the way for more unified and controllable video learning frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12761
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Moaw: Unleashing Motion Awareness for Video Diffusion Models
Zhang, Tianqi
Wang, Ziyi
Zheng, Wenzhao
Chen, Weiliang
Huang, Yuanhui
Huang, Zhengyang
Zhou, Jie
Lu, Jiwen
Computer Vision and Pattern Recognition
Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot setting. Motivated by these findings, we investigate whether supervised training can more fully harness the tracking capability of video diffusion models. To this end, we propose Moaw, a framework that unleashes motion awareness for video diffusion models and leverages it to facilitate motion transfer. Specifically, we train a diffusion model for motion perception, shifting its modality from image-to-video generation to video-to-dense-tracking. We then construct a motion-labeled dataset to identify features that encode the strongest motion information, and inject them into a structurally identical video generation model. Owing to the homogeneity between the two networks, these features can be naturally adapted in a zero-shot manner, enabling motion transfer without additional adapters. Our work provides a new paradigm for bridging generative modeling and motion understanding, paving the way for more unified and controllable video learning frameworks.
title Moaw: Unleashing Motion Awareness for Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.12761