Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Yan-Bo, Lin, Kevin, Yang, Zhengyuan, Li, Linjie, Wang, Jianfeng, Lin, Chung-Ching, Wang, Xiaofei, Bertasius, Gedas, Wang, Lijuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915214281146368
author Lin, Yan-Bo
Lin, Kevin
Yang, Zhengyuan
Li, Linjie
Wang, Jianfeng
Lin, Chung-Ching
Wang, Xiaofei
Bertasius, Gedas
Wang, Lijuan
author_facet Lin, Yan-Bo
Lin, Kevin
Yang, Zhengyuan
Li, Linjie
Wang, Jianfeng
Lin, Chung-Ching
Wang, Xiaofei
Bertasius, Gedas
Wang, Lijuan
contents In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AvED-Bench, designed explicitly for zero-shot audio-video editing. AvED-Bench includes 110 videos, each with a 10-second duration, spanning 11 categories from VGGSound. It offers diverse prompts and scenarios that require precise alignment between auditory and visual elements, enabling robust evaluation. We identify limitations in existing zero-shot audio and video editing methods, particularly in synchronization and coherence between modalities, which often result in inconsistent outcomes. To address these challenges, we propose AvED, a zero-shot cross-modal delta denoising framework that leverages audio-video interactions to achieve synchronized and coherent edits. AvED demonstrates superior results on both AvED-Bench and the recent OAVE dataset to validate its generalization capabilities. Results are available at https://genjib.github.io/project_page/AVED/index.html
format Preprint
id arxiv_https___arxiv_org_abs_2503_20782
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
Lin, Yan-Bo
Lin, Kevin
Yang, Zhengyuan
Li, Linjie
Wang, Jianfeng
Lin, Chung-Ching
Wang, Xiaofei
Bertasius, Gedas
Wang, Lijuan
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Sound
Audio and Speech Processing
In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AvED-Bench, designed explicitly for zero-shot audio-video editing. AvED-Bench includes 110 videos, each with a 10-second duration, spanning 11 categories from VGGSound. It offers diverse prompts and scenarios that require precise alignment between auditory and visual elements, enabling robust evaluation. We identify limitations in existing zero-shot audio and video editing methods, particularly in synchronization and coherence between modalities, which often result in inconsistent outcomes. To address these challenges, we propose AvED, a zero-shot cross-modal delta denoising framework that leverages audio-video interactions to achieve synchronized and coherent edits. AvED demonstrates superior results on both AvED-Bench and the recent OAVE dataset to validate its generalization capabilities. Results are available at https://genjib.github.io/project_page/AVED/index.html
title Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.20782