MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tong, Alcazar, Juan C Leon, Escorcia, Victor, Ghanem, Bernard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918245114576896
author Zhang, Tong
Alcazar, Juan C Leon
Escorcia, Victor
Ghanem, Bernard
author_facet Zhang, Tong
Alcazar, Juan C Leon
Escorcia, Victor
Ghanem, Bernard
contents We present MoCA-Video, a training-free framework for semantic mixing in videos. Operating in the latent space of a frozen video diffusion model, MoCA-Video utilizes class-agnostic segmentation with diagonal denoising scheduler to localize and track the target object across frames. To ensure temporal stability under semantic shifts, we introduce momentum-based correction to approximate novel hybrid distributions beyond trained data distribution, alongside a light gamma residual module that smooths out visual artifacts. We evaluate model's performance using SSIM, LPIPS, and a proposed metric, \metricnameabbr, which quantifies semantic alignment between reference and output. Extensive evaluation demonstrates that our model consistently outperforms both training-free and trained baselines, achieving superior semantic mixing and temporal coherence without retraining. Results establish that structured manipulation of diffusion noise trajectories enables controllable and high-quality video editing under semantic shifts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
Zhang, Tong
Alcazar, Juan C Leon
Escorcia, Victor
Ghanem, Bernard
Computer Vision and Pattern Recognition
Artificial Intelligence
We present MoCA-Video, a training-free framework for semantic mixing in videos. Operating in the latent space of a frozen video diffusion model, MoCA-Video utilizes class-agnostic segmentation with diagonal denoising scheduler to localize and track the target object across frames. To ensure temporal stability under semantic shifts, we introduce momentum-based correction to approximate novel hybrid distributions beyond trained data distribution, alongside a light gamma residual module that smooths out visual artifacts. We evaluate model's performance using SSIM, LPIPS, and a proposed metric, \metricnameabbr, which quantifies semantic alignment between reference and output. Extensive evaluation demonstrates that our model consistently outperforms both training-free and trained baselines, achieving superior semantic mixing and temporal coherence without retraining. Results establish that structured manipulation of diffusion noise trajectories enables controllable and high-quality video editing under semantic shifts.
title MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.01004