InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Haojie, Yang, Yixin, Yang, Siqi, Weng, Shuchen, Shi, Boxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918509320077312
author Zheng, Haojie
Yang, Yixin
Yang, Siqi
Weng, Shuchen
Shi, Boxin
author_facet Zheng, Haojie
Yang, Yixin
Yang, Siqi
Weng, Shuchen
Shi, Boxin
contents Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose InstructAV2AV, the first end-to-end framework for instruction-guided audio-video joint editing. We first develop a scalable data synthesis pipeline and construct InsAVE-80K, the first large-scale audio-video editing dataset with high-quality source-to-target pairs. With this data foundation, we adapt an audio-video generation backbone to leverage its robust priors. We concatenate the audio-video input with noisy latent codes to anchor the source context, propose the source-instruction gated attention to improve instruction following and content preservation, and introduce a two-stage training strategy to effectively transfer these pre-trained priors. Extensive experiments demonstrate that InstructAV2AV outperforms state-of-the-art methods across 11 metrics spanning three aspects on two evaluation sets, highlighting its potential for controllable content creation. Project page: https://hjzheng.net/projects/InstructAV2AV/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18467
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
Zheng, Haojie
Yang, Yixin
Yang, Siqi
Weng, Shuchen
Shi, Boxin
Computer Vision and Pattern Recognition
Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose InstructAV2AV, the first end-to-end framework for instruction-guided audio-video joint editing. We first develop a scalable data synthesis pipeline and construct InsAVE-80K, the first large-scale audio-video editing dataset with high-quality source-to-target pairs. With this data foundation, we adapt an audio-video generation backbone to leverage its robust priors. We concatenate the audio-video input with noisy latent codes to anchor the source context, propose the source-instruction gated attention to improve instruction following and content preservation, and introduce a two-stage training strategy to effectively transfer these pre-trained priors. Extensive experiments demonstrate that InstructAV2AV outperforms state-of-the-art methods across 11 metrics spanning three aspects on two evaluation sets, highlighting its potential for controllable content creation. Project page: https://hjzheng.net/projects/InstructAV2AV/.
title InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.18467