Object-AVEdit: An Object-level Audio-Visual Editing Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fu, Youquan, Si, Ruiyang, Wang, Hongfa, Zhou, Dongzhan, Sun, Jiacheng, Luo, Ping, Hu, Di, Zhang, Hongyuan, Li, Xuelong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915527374405632
author Fu, Youquan
Si, Ruiyang
Wang, Hongfa
Zhou, Dongzhan
Sun, Jiacheng
Luo, Ping
Hu, Di
Zhang, Hongyuan
Li, Xuelong
author_facet Fu, Youquan
Si, Ruiyang
Wang, Hongfa
Zhou, Dongzhan
Sun, Jiacheng
Luo, Ping
Hu, Di
Zhang, Hongyuan
Li, Xuelong
contents There is a high demand for audio-visual editing in video post-production and the film making field. While numerous models have explored audio and video editing, they struggle with object-level audio-visual operations. Specifically, object-level audio-visual editing requires the ability to perform object addition, replacement, and removal across both audio and visual modalities, while preserving the structural information of the source instances during the editing process. In this paper, we present \textbf{Object-AVEdit}, achieving the object-level audio-visual editing based on the inversion-regeneration paradigm. To achieve the object-level controllability during editing, we develop a word-to-sounding-object well-aligned audio generation model, bridging the gap in object-controllability between audio and current video generation models. Meanwhile, to achieve the better structural information preservation and object-level editing effect, we propose an inversion-regeneration holistically-optimized editing algorithm, ensuring both information retention during the inversion and better regeneration effect. Extensive experiments demonstrate that our editing model achieved advanced results in both audio-video object-level editing tasks with fine audio-visual semantic alignment. In addition, our developed audio generation model also achieved advanced performance. More results on our project page: https://gewu-lab.github.io/Object_AVEdit-website/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00050
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Object-AVEdit: An Object-level Audio-Visual Editing Model
Fu, Youquan
Si, Ruiyang
Wang, Hongfa
Zhou, Dongzhan
Sun, Jiacheng
Luo, Ping
Hu, Di
Zhang, Hongyuan
Li, Xuelong
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
There is a high demand for audio-visual editing in video post-production and the film making field. While numerous models have explored audio and video editing, they struggle with object-level audio-visual operations. Specifically, object-level audio-visual editing requires the ability to perform object addition, replacement, and removal across both audio and visual modalities, while preserving the structural information of the source instances during the editing process. In this paper, we present \textbf{Object-AVEdit}, achieving the object-level audio-visual editing based on the inversion-regeneration paradigm. To achieve the object-level controllability during editing, we develop a word-to-sounding-object well-aligned audio generation model, bridging the gap in object-controllability between audio and current video generation models. Meanwhile, to achieve the better structural information preservation and object-level editing effect, we propose an inversion-regeneration holistically-optimized editing algorithm, ensuring both information retention during the inversion and better regeneration effect. Extensive experiments demonstrate that our editing model achieved advanced results in both audio-video object-level editing tasks with fine audio-visual semantic alignment. In addition, our developed audio generation model also achieved advanced performance. More results on our project page: https://gewu-lab.github.io/Object_AVEdit-website/.
title Object-AVEdit: An Object-level Audio-Visual Editing Model
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.00050