MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yen-Siang, Huang, Chi-Pin, Yang, Fu-En, Wang, Yu-Chiang Frank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915160352882688
author Wu, Yen-Siang
Huang, Chi-Pin
Yang, Fu-En
Wang, Yu-Chiang Frank
author_facet Wu, Yen-Siang
Huang, Chi-Pin
Yang, Fu-En
Wang, Yu-Chiang Frank
contents Text-to-video (T2V) diffusion models have shown promising capabilities in synthesizing realistic videos from input text prompts. However, the input text description alone provides limited control over the precise objects movements and camera framing. In this work, we tackle the motion customization problem, where a reference video is provided as motion guidance. While most existing methods choose to fine-tune pre-trained diffusion models to reconstruct the frame differences of the reference video, we observe that such strategy suffer from content leakage from the reference video, and they cannot capture complex motion accurately. To address this issue, we propose MotionMatcher, a motion customization framework that fine-tunes the pre-trained T2V diffusion model at the feature level. Instead of using pixel-level objectives, MotionMatcher compares high-level, spatio-temporal motion features to fine-tune diffusion models, ensuring precise motion learning. For the sake of memory efficiency and accessibility, we utilize a pre-trained T2V diffusion model, which contains considerable prior knowledge about video motion, to compute these motion features. In our experiments, we demonstrate state-of-the-art motion customization performances, validating the design of our framework.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
Wu, Yen-Siang
Huang, Chi-Pin
Yang, Fu-En
Wang, Yu-Chiang Frank
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Text-to-video (T2V) diffusion models have shown promising capabilities in synthesizing realistic videos from input text prompts. However, the input text description alone provides limited control over the precise objects movements and camera framing. In this work, we tackle the motion customization problem, where a reference video is provided as motion guidance. While most existing methods choose to fine-tune pre-trained diffusion models to reconstruct the frame differences of the reference video, we observe that such strategy suffer from content leakage from the reference video, and they cannot capture complex motion accurately. To address this issue, we propose MotionMatcher, a motion customization framework that fine-tunes the pre-trained T2V diffusion model at the feature level. Instead of using pixel-level objectives, MotionMatcher compares high-level, spatio-temporal motion features to fine-tune diffusion models, ensuring precise motion learning. For the sake of memory efficiency and accessibility, we utilize a pre-trained T2V diffusion model, which contains considerable prior knowledge about video motion, to compute these motion features. In our experiments, we demonstrate state-of-the-art motion customization performances, validating the design of our framework.
title MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.13234