MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Yanchen, Sun, Yanan, Xing, Zhening, Gao, Junyao, Chen, Kai, Pei, Wenjie
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911069761437696
author Liu, Yanchen
Sun, Yanan
Xing, Zhening
Gao, Junyao
Chen, Kai
Pei, Wenjie
author_facet Liu, Yanchen
Sun, Yanan
Xing, Zhening
Gao, Junyao
Chen, Kai
Pei, Wenjie
contents Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a training-free framework capable of parsing reference-target correspondences in a fine-grained manner, thereby achieving high-fidelity motion transfer while preserving coherence in appearance. To be specific, MotionShot first performs semantic feature matching to ensure high-level alignments between the reference and target objects. It then further establishes low-level morphological alignments through reference-to-target shape retargeting. By encoding motion with temporal attention, our MotionShot can coherently transfer motion across objects, even in the presence of significant appearance and structure disparities, demonstrated by extensive experiments. The project page is available at: https://motionshot.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
Liu, Yanchen
Sun, Yanan
Xing, Zhening
Gao, Junyao
Chen, Kai
Pei, Wenjie
Computer Vision and Pattern Recognition
Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a training-free framework capable of parsing reference-target correspondences in a fine-grained manner, thereby achieving high-fidelity motion transfer while preserving coherence in appearance. To be specific, MotionShot first performs semantic feature matching to ensure high-level alignments between the reference and target objects. It then further establishes low-level morphological alignments through reference-to-target shape retargeting. By encoding motion with temporal attention, our MotionShot can coherently transfer motion across objects, even in the presence of significant appearance and structure disparities, demonstrated by extensive experiments. The project page is available at: https://motionshot.github.io/.
title MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.16310