MotionV2V: Editing Motion in a Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burgert, Ryan, Herrmann, Charles, Cole, Forrester, Ryoo, Michael S, Wadhwa, Neal, Voynov, Andrey, Ruiz, Nataniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912729086820352
author Burgert, Ryan
Herrmann, Charles
Cole, Forrester
Ryoo, Michael S
Wadhwa, Neal
Voynov, Andrey
Ruiz, Nataniel
author_facet Burgert, Ryan
Herrmann, Charles
Cole, Forrester
Ryoo, Michael S
Wadhwa, Neal
Voynov, Andrey
Ruiz, Nataniel
contents While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has explored motion controllability as a means to enhance text-to-video generation or image animation; however, we identify precise motion control as a promising yet under-explored paradigm for editing existing videos. In this work, we propose modifying video motion by directly editing sparse trajectories extracted from the input. We term the deviation between input and output trajectories a "motion edit" and demonstrate that this representation, when coupled with a generative backbone, enables powerful video editing capabilities. To achieve this, we introduce a pipeline for generating "motion counterfactuals", video pairs that share identical content but distinct motion, and we fine-tune a motion-conditioned video diffusion architecture on this dataset. Our approach allows for edits that start at any timestamp and propagate naturally. In a four-way head-to-head user study, our model achieves over 65 percent preference against prior work. Please see our project page: https://ryanndagreat.github.io/MotionV2V
format Preprint
id arxiv_https___arxiv_org_abs_2511_20640
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MotionV2V: Editing Motion in a Video
Burgert, Ryan
Herrmann, Charles
Cole, Forrester
Ryoo, Michael S
Wadhwa, Neal
Voynov, Andrey
Ruiz, Nataniel
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has explored motion controllability as a means to enhance text-to-video generation or image animation; however, we identify precise motion control as a promising yet under-explored paradigm for editing existing videos. In this work, we propose modifying video motion by directly editing sparse trajectories extracted from the input. We term the deviation between input and output trajectories a "motion edit" and demonstrate that this representation, when coupled with a generative backbone, enables powerful video editing capabilities. To achieve this, we introduce a pipeline for generating "motion counterfactuals", video pairs that share identical content but distinct motion, and we fine-tune a motion-conditioned video diffusion architecture on this dataset. Our approach allows for edits that start at any timestamp and propagate naturally. In a four-way head-to-head user study, our model achieves over 65 percent preference against prior work. Please see our project page: https://ryanndagreat.github.io/MotionV2V
title MotionV2V: Editing Motion in a Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2511.20640