Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choi, Saemee, Jeong, Sohyun, Jang, Hyojin, Choo, Jaegul, Kim, Jinhee
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914211481780224
author Choi, Saemee
Jeong, Sohyun
Jang, Hyojin
Choo, Jaegul
Kim, Jinhee
author_facet Choi, Saemee
Jeong, Sohyun
Jang, Hyojin
Choo, Jaegul
Kim, Jinhee
contents We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces $ρ$-start sampling and dilated dual masking to construct structured noise maps that enable coherent and accurate edits. To further enhance visual fidelity, we present zero image guidance, a controllable negative prompt strategy. Extensive experiments demonstrate that VINO faithfully incorporates the reference image into video edits, achieving strong performance compared to state-of-the-art baselines, all without any test-time or instance-specific training.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
Choi, Saemee
Jeong, Sohyun
Jang, Hyojin
Choo, Jaegul
Kim, Jinhee
Computer Vision and Pattern Recognition
We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces $ρ$-start sampling and dilated dual masking to construct structured noise maps that enable coherent and accurate edits. To further enhance visual fidelity, we present zero image guidance, a controllable negative prompt strategy. Extensive experiments demonstrate that VINO faithfully incorporates the reference image into video edits, achieving strong performance compared to state-of-the-art baselines, all without any test-time or instance-specific training.
title Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.12520