Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Yue, Liu, Yulong, Zhu, Qiyuan, Yang, Ayden, Feng, Kunyu, Zhang, Xinhua, Yan, Zexuan, Li, Zhifeng, Han, Sirui, Qi, Chenyang, Chen, Qifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917367107289088
author Ma, Yue
Liu, Yulong
Zhu, Qiyuan
Yang, Ayden
Feng, Kunyu
Zhang, Xinhua
Yan, Zexuan
Li, Zhifeng
Han, Sirui
Qi, Chenyang
Chen, Qifeng
author_facet Ma, Yue
Liu, Yulong
Zhu, Qiyuan
Yang, Ayden
Feng, Kunyu
Zhang, Xinhua
Yan, Zexuan
Li, Zhifeng
Han, Sirui
Qi, Chenyang
Chen, Qifeng
contents Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoRAs) finetuning to obtain better performance. However, existing adaptation-based motion transfer still suffers from motion inconsistency and tuning inefficiency when applied to large video diffusion transformers. Naive two-stage LoRA tuning struggles to maintain motion consistency between generated and input videos due to the inherent spatial-temporal coupling in the 3D attention operator. Additionally, they require time-consuming fine-tuning processes in both stages. To tackle these issues, we propose Follow-Your-Motion, an efficient two-stage video motion transfer framework that finetunes a powerful video diffusion transformer to synthesize complex motion. Specifically, we propose a spatial-temporal decoupled LoRA to decouple the attention architecture for spatial appearance and temporal motion processing. During the second training stage, we design the sparse motion sampling and adaptive RoPE to accelerate the tuning speed. To address the lack of a benchmark for this field, we introduce MotionBench, a comprehensive benchmark comprising diverse motion, including creative camera motion, single object motion, multiple object motion, and complex human motion. We show extensive evaluations on MotionBench to verify the superiority of Follow-Your-Motion.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
Ma, Yue
Liu, Yulong
Zhu, Qiyuan
Yang, Ayden
Feng, Kunyu
Zhang, Xinhua
Yan, Zexuan
Li, Zhifeng
Han, Sirui
Qi, Chenyang
Chen, Qifeng
Computer Vision and Pattern Recognition
Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoRAs) finetuning to obtain better performance. However, existing adaptation-based motion transfer still suffers from motion inconsistency and tuning inefficiency when applied to large video diffusion transformers. Naive two-stage LoRA tuning struggles to maintain motion consistency between generated and input videos due to the inherent spatial-temporal coupling in the 3D attention operator. Additionally, they require time-consuming fine-tuning processes in both stages. To tackle these issues, we propose Follow-Your-Motion, an efficient two-stage video motion transfer framework that finetunes a powerful video diffusion transformer to synthesize complex motion. Specifically, we propose a spatial-temporal decoupled LoRA to decouple the attention architecture for spatial appearance and temporal motion processing. During the second training stage, we design the sparse motion sampling and adaptive RoPE to accelerate the tuning speed. To address the lack of a benchmark for this field, we introduce MotionBench, a comprehensive benchmark comprising diverse motion, including creative camera motion, single object motion, multiple object motion, and complex human motion. We show extensive evaluations on MotionBench to verify the superiority of Follow-Your-Motion.
title Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05207