Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Qingyu, Wu, Jianzong, Bai, Jinbin, Zhang, Jiangning, Qi, Lu, Tong, Yunhai, Li, Xiangtai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912520339456000
author Shi, Qingyu
Wu, Jianzong
Bai, Jinbin
Zhang, Jiangning
Qi, Lu
Tong, Yunhai
Li, Xiangtai
author_facet Shi, Qingyu
Wu, Jianzong
Bai, Jinbin
Zhang, Jiangning
Qi, Lu
Tong, Yunhai
Li, Xiangtai
contents The motion transfer task aims to transfer motion from a source video to newly generated videos, requiring the model to decouple motion from appearance. Previous diffusion-based methods primarily rely on separate spatial and temporal attention mechanisms within the 3D U-Net. In contrast, state-of-the-art video Diffusion Transformers (DiT) models use 3D full attention, which does not explicitly separate temporal and spatial information. Thus, the interaction between spatial and temporal dimensions makes decoupling motion and appearance more challenging for DiT models. In this paper, we propose DeT, a method that adapts DiT models to improve motion transfer ability. Our approach introduces a simple yet effective temporal kernel to smooth DiT features along the temporal dimension, facilitating the decoupling of foreground motion from background appearance. Meanwhile, the temporal kernel effectively captures temporal variations in DiT features, which are closely related to motion. Moreover, we introduce explicit supervision along dense trajectories in the latent feature space to further enhance motion consistency. Additionally, we present MTBench, a general and challenging benchmark for motion transfer. We also introduce a hybrid motion fidelity metric that considers both the global and local motion similarity. Therefore, our work provides a more comprehensive evaluation than previous works. Extensive experiments on MTBench demonstrate that DeT achieves the best trade-off between motion fidelity and edit fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
Shi, Qingyu
Wu, Jianzong
Bai, Jinbin
Zhang, Jiangning
Qi, Lu
Tong, Yunhai
Li, Xiangtai
Computer Vision and Pattern Recognition
The motion transfer task aims to transfer motion from a source video to newly generated videos, requiring the model to decouple motion from appearance. Previous diffusion-based methods primarily rely on separate spatial and temporal attention mechanisms within the 3D U-Net. In contrast, state-of-the-art video Diffusion Transformers (DiT) models use 3D full attention, which does not explicitly separate temporal and spatial information. Thus, the interaction between spatial and temporal dimensions makes decoupling motion and appearance more challenging for DiT models. In this paper, we propose DeT, a method that adapts DiT models to improve motion transfer ability. Our approach introduces a simple yet effective temporal kernel to smooth DiT features along the temporal dimension, facilitating the decoupling of foreground motion from background appearance. Meanwhile, the temporal kernel effectively captures temporal variations in DiT features, which are closely related to motion. Moreover, we introduce explicit supervision along dense trajectories in the latent feature space to further enhance motion consistency. Additionally, we present MTBench, a general and challenging benchmark for motion transfer. We also introduce a hybrid motion fidelity metric that considers both the global and local motion similarity. Therefore, our work provides a more comprehensive evaluation than previous works. Extensive experiments on MTBench demonstrate that DeT achieves the best trade-off between motion fidelity and edit fidelity.
title Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.17350