AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Xinyue, Yang, Xiaoran, Zhang, Lipan, Yang, Jianxuan, Wang, Zhao, Luan, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909925646532608
author Guo, Xinyue
Yang, Xiaoran
Zhang, Lipan
Yang, Jianxuan
Wang, Zhao
Luan, Jian
author_facet Guo, Xinyue
Yang, Xiaoran
Zhang, Lipan
Yang, Jianxuan
Wang, Zhao
Luan, Jian
contents Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
Guo, Xinyue
Yang, Xiaoran
Zhang, Lipan
Yang, Jianxuan
Wang, Zhao
Luan, Jian
Multimedia
Computer Vision and Pattern Recognition
Sound
Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.
title AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
topic Multimedia
Computer Vision and Pattern Recognition
Sound
url https://arxiv.org/abs/2511.21146