VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Kam Man, Tian, Zeyue, Ji, Liya, Chen, Qifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918145037434880
author Wu, Kam Man
Tian, Zeyue
Ji, Liya
Chen, Qifeng
author_facet Wu, Kam Man
Tian, Zeyue
Ji, Liya
Chen, Qifeng
contents Video and audio inpainting for mixed audio-visual content has become a crucial task in multimedia editing recently. However, precisely removing an object and its corresponding audio from a video without affecting the rest of the scene remains a significant challenge. To address this, we propose VAInpaint, a novel pipeline that first utilizes a segmentation model to generate masks and guide a video inpainting model in removing objects. At the same time, an LLM then analyzes the scene globally, while a region-specific model provides localized descriptions. Both the overall and regional descriptions will be inputted into an LLM, which will refine the content and turn it into text queries for our text-driven audio separation model. Our audio separation model is fine-tuned on a customized dataset comprising segmented MUSIC instrument images and VGGSound backgrounds to enhance its generalization performance. Experiments show that our method achieves performance comparable to current benchmarks in both audio and video inpainting.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
Wu, Kam Man
Tian, Zeyue
Ji, Liya
Chen, Qifeng
Multimedia
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Video and audio inpainting for mixed audio-visual content has become a crucial task in multimedia editing recently. However, precisely removing an object and its corresponding audio from a video without affecting the rest of the scene remains a significant challenge. To address this, we propose VAInpaint, a novel pipeline that first utilizes a segmentation model to generate masks and guide a video inpainting model in removing objects. At the same time, an LLM then analyzes the scene globally, while a region-specific model provides localized descriptions. Both the overall and regional descriptions will be inputted into an LLM, which will refine the content and turn it into text queries for our text-driven audio separation model. Our audio separation model is fine-tuned on a customized dataset comprising segmented MUSIC instrument images and VGGSound backgrounds to enhance its generalization performance. Experiments show that our method achieves performance comparable to current benchmarks in both audio and video inpainting.
title VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
topic Multimedia
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.17022