Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zuo, Yi, Li, Lingling, Jiao, Licheng, Liu, Fang, Liu, Xu, Ma, Wenping, Yang, Shuyuan, Guo, Yuwei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912072925708288
author Zuo, Yi
Li, Lingling
Jiao, Licheng
Liu, Fang
Liu, Xu
Ma, Wenping
Yang, Shuyuan
Guo, Yuwei
author_facet Zuo, Yi
Li, Lingling
Jiao, Licheng
Liu, Fang
Liu, Xu
Ma, Wenping
Yang, Shuyuan
Guo, Yuwei
contents Existing diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce Edit-Your-Motion, a video motion editing method that tackles these challenges through one-shot fine-tuning on unseen cases. Specifically, firstly, we utilized DDIM inversion to initialize the noise, preserving the appearance of the source video and designed a lightweight motion attention adapter module to enhance motion fidelity. DDIM inversion aims to obtain the implicit representations by estimating the prediction noise from the source video, which serves as a starting point for the sampling process, ensuring the appearance consistency between the source and edited videos. The Motion Attention Module (MA) enhances the model's motion editing ability by resolving the conflict between the skeleton features and the appearance features. Secondly, to effectively decouple motion and appearance of source video, we design a spatio-temporal two-stage learning strategy (STL). In the first stage, we focus on learning temporal features of human motion and propose recurrent causal attention (RCA) to ensure consistency between video frames. In the second stage, we shift focus on learning the appearance features of the source video. With Edit-Your-Motion, users can edit the motion of humans in the source video, creating more engaging and diverse content. Extensive qualitative and quantitative experiments, along with user preference studies, show that Edit-Your-Motion outperforms other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04496
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
Zuo, Yi
Li, Lingling
Jiao, Licheng
Liu, Fang
Liu, Xu
Ma, Wenping
Yang, Shuyuan
Guo, Yuwei
Computer Vision and Pattern Recognition
Existing diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce Edit-Your-Motion, a video motion editing method that tackles these challenges through one-shot fine-tuning on unseen cases. Specifically, firstly, we utilized DDIM inversion to initialize the noise, preserving the appearance of the source video and designed a lightweight motion attention adapter module to enhance motion fidelity. DDIM inversion aims to obtain the implicit representations by estimating the prediction noise from the source video, which serves as a starting point for the sampling process, ensuring the appearance consistency between the source and edited videos. The Motion Attention Module (MA) enhances the model's motion editing ability by resolving the conflict between the skeleton features and the appearance features. Secondly, to effectively decouple motion and appearance of source video, we design a spatio-temporal two-stage learning strategy (STL). In the first stage, we focus on learning temporal features of human motion and propose recurrent causal attention (RCA) to ensure consistency between video frames. In the second stage, we shift focus on learning the appearance features of the source video. With Edit-Your-Motion, users can edit the motion of humans in the source video, creating more engaging and diverse content. Extensive qualitative and quantitative experiments, along with user preference studies, show that Edit-Your-Motion outperforms other methods.
title Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.04496