Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Song, Chen, Xinyu, Wang, Qian, Li, Liang, Yi, Zili, Feng, Junlan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911687678885888
author Wu, Song
Chen, Xinyu
Wang, Qian
Li, Liang
Yi, Zili
Feng, Junlan
author_facet Wu, Song
Chen, Xinyu
Wang, Qian
Li, Liang
Yi, Zili
Feng, Junlan
contents Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the rich information embedded within noisy latent, leading to unsatisfactory results. To address this, we propose a \textit{tuning-free, instruction-based} video editing framework. We approach video editing from the perspective of noisy latent: we design a Structural Noise Initialization Strategy (SNIS) to secure a superior editing starting point by assigning higher noise levels to edited regions (to facilitate content change) and lower noise levels to unedited regions (to maintain content consistency). We introduce a Noise Guidance Mechanism (NGM), which leverages the video prior in the generative model and effectively integrates rich information within the noisy latent to guide the denoising process, thereby preserving unedited content and overall visual coherence. Experiments show that our proposed method achieves better visual quality and state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15533
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
Wu, Song
Chen, Xinyu
Wang, Qian
Li, Liang
Yi, Zili
Feng, Junlan
Computer Vision and Pattern Recognition
Artificial Intelligence
Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the rich information embedded within noisy latent, leading to unsatisfactory results. To address this, we propose a \textit{tuning-free, instruction-based} video editing framework. We approach video editing from the perspective of noisy latent: we design a Structural Noise Initialization Strategy (SNIS) to secure a superior editing starting point by assigning higher noise levels to edited regions (to facilitate content change) and lower noise levels to unedited regions (to maintain content consistency). We introduce a Noise Guidance Mechanism (NGM), which leverages the video prior in the generative model and effectively integrates rich information within the noisy latent to guide the denoising process, thereby preserving unedited content and overall visual coherence. Experiments show that our proposed method achieves better visual quality and state-of-the-art performance.
title Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.15533