Audio-Guided Visual Editing with Complex Multi-Modal Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Hyeonyu, Jeong, Seokhoon, Han, Seonghee, Choi, Chanhyuk, Kim, Taehwan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912557636255744
author Kim, Hyeonyu
Jeong, Seokhoon
Han, Seonghee
Choi, Chanhyuk
Kim, Taehwan
author_facet Kim, Hyeonyu
Jeong, Seokhoon
Han, Seonghee
Choi, Chanhyuk
Kim, Taehwan
contents Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlighting the need for additional non-text editing prompts. In this work, we introduce a novel audio-guided visual editing framework that can handle complex editing tasks with multiple text and audio prompts without requiring additional training. Existing audio-guided visual editing methods often necessitate training on specific datasets to align audio with text, limiting their generalization to real-world situations. We leverage a pre-trained multi-modal encoder with strong zero-shot capabilities and integrate diverse audio into visual editing tasks, by alleviating the discrepancy between the audio encoder space and the diffusion model's prompt encoder space. Additionally, we propose a novel approach to handle complex scenarios with multiple and multi-modal editing prompts through our separate noise branching and adaptive patch selection. Our comprehensive experiments on diverse editing tasks demonstrate that our framework excels in handling complicated editing scenarios by incorporating rich information from audio, where text-only approaches fail.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio-Guided Visual Editing with Complex Multi-Modal Prompts
Kim, Hyeonyu
Jeong, Seokhoon
Han, Seonghee
Choi, Chanhyuk
Kim, Taehwan
Computer Vision and Pattern Recognition
Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlighting the need for additional non-text editing prompts. In this work, we introduce a novel audio-guided visual editing framework that can handle complex editing tasks with multiple text and audio prompts without requiring additional training. Existing audio-guided visual editing methods often necessitate training on specific datasets to align audio with text, limiting their generalization to real-world situations. We leverage a pre-trained multi-modal encoder with strong zero-shot capabilities and integrate diverse audio into visual editing tasks, by alleviating the discrepancy between the audio encoder space and the diffusion model's prompt encoder space. Additionally, we propose a novel approach to handle complex scenarios with multiple and multi-modal editing prompts through our separate noise branching and adaptive patch selection. Our comprehensive experiments on diverse editing tasks demonstrate that our framework excels in handling complicated editing scenarios by incorporating rich information from audio, where text-only approaches fail.
title Audio-Guided Visual Editing with Complex Multi-Modal Prompts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.20379