Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Weihan, Cheng, Kan Jen, Saito, Koichi, Mirza, Muhammad Jehanzeb, Li, Tingle, Liu, Yisi, Liu, Alexander H., Wang, Liming, Ishii, Masato, Shibuya, Takashi, Mitsufuji, Yuki, Anumanchipalli, Gopala, Liang, Paul Pu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911318787751936
author Xu, Weihan
Cheng, Kan Jen
Saito, Koichi
Mirza, Muhammad Jehanzeb
Li, Tingle
Liu, Yisi
Liu, Alexander H.
Wang, Liming
Ishii, Masato
Shibuya, Takashi
Mitsufuji, Yuki
Anumanchipalli, Gopala
Liang, Paul Pu
author_facet Xu, Weihan
Cheng, Kan Jen
Saito, Koichi
Mirza, Muhammad Jehanzeb
Li, Tingle
Liu, Yisi
Liu, Alexander H.
Wang, Liming
Ishii, Masato
Shibuya, Takashi
Mitsufuji, Yuki
Anumanchipalli, Gopala
Liang, Paul Pu
contents Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual data before and after targeted edits, and the heterogeneity across modalities. To address the data and modeling challenges in joint audio-visual editing, we introduce SAVEBench, a paired audiovisual dataset with text and mask conditions to enable object-grounded source-to-target learning. With SAVEBench, we train the Schrodinger Audio-Visual Editor (SAVE), an end-to-end flow-matching model that edits audio and video in parallel while keeping them aligned throughout processing. SAVE incorporates a Schrodinger Bridge that learns a direct transport from source to target audiovisual mixtures. Our evaluation demonstrates that the proposed SAVE model is able to remove the target objects in audio and visual content while preserving the remaining content, with stronger temporal synchronization and audiovisual semantic correspondence compared with pairwise combinations of an audio editor and a video editor.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
Xu, Weihan
Cheng, Kan Jen
Saito, Koichi
Mirza, Muhammad Jehanzeb
Li, Tingle
Liu, Yisi
Liu, Alexander H.
Wang, Liming
Ishii, Masato
Shibuya, Takashi
Mitsufuji, Yuki
Anumanchipalli, Gopala
Liang, Paul Pu
Computer Vision and Pattern Recognition
Multimedia
Sound
Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual data before and after targeted edits, and the heterogeneity across modalities. To address the data and modeling challenges in joint audio-visual editing, we introduce SAVEBench, a paired audiovisual dataset with text and mask conditions to enable object-grounded source-to-target learning. With SAVEBench, we train the Schrodinger Audio-Visual Editor (SAVE), an end-to-end flow-matching model that edits audio and video in parallel while keeping them aligned throughout processing. SAVE incorporates a Schrodinger Bridge that learns a direct transport from source to target audiovisual mixtures. Our evaluation demonstrates that the proposed SAVE model is able to remove the target objects in audio and visual content while preserving the remaining content, with stronger temporal synchronization and audiovisual semantic correspondence compared with pairwise combinations of an audio editor and a video editor.
title Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
topic Computer Vision and Pattern Recognition
Multimedia
Sound
url https://arxiv.org/abs/2512.12875