InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Haojie, Yang, Yixin, Yang, Siqi, Weng, Shuchen, Shi, Boxin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
by: Zheng, Haojie, et al.
Published: (2025)
by: Zheng, Haojie, et al.
Published: (2025)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
by: Weng, Shuchen, et al.
Published: (2024)
by: Weng, Shuchen, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
by: Guo, Xinyue, et al.
Published: (2025)
by: Guo, Xinyue, et al.
Published: (2025)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
by: Liu, Tengfei, et al.
Published: (2026)
by: Liu, Tengfei, et al.
Published: (2026)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
by: Cai, Ziqi, et al.
Published: (2026)
by: Cai, Ziqi, et al.
Published: (2026)
L-C4: Language-Based Video Colorization for Creative and Consistent Color
by: Chang, Zheng, et al.
Published: (2024)
by: Chang, Zheng, et al.
Published: (2024)
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
by: Cao, Zhe, et al.
Published: (2025)
by: Cao, Zhe, et al.
Published: (2025)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
by: Ye, Zhen, et al.
Published: (2026)
by: Ye, Zhen, et al.
Published: (2026)
Affective Image Editing: Shaping Emotional Factors via Text Descriptions
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
by: Xia, Shuhan, et al.
Published: (2025)
by: Xia, Shuhan, et al.
Published: (2025)
InstructVEdit: A Holistic Approach for Instructional Video Editing
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
Language-guided Image Reflection Separation
by: Zhong, Haofeng, et al.
Published: (2024)
by: Zhong, Haofeng, et al.
Published: (2024)
Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
E2VIDiff: Perceptual Events-to-Video Reconstruction using Diffusion Priors
by: Liang, Jinxiu, et al.
Published: (2024)
by: Liang, Jinxiu, et al.
Published: (2024)
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
by: Zhang, Peixuan, et al.
Published: (2026)
by: Zhang, Peixuan, et al.
Published: (2026)
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
by: Li, Guangyao, et al.
Published: (2026)
by: Li, Guangyao, et al.
Published: (2026)
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025)
by: Chatziagapi, Aggelina, et al.
Published: (2025)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
by: Feng, Hengyi, et al.
Published: (2026)
by: Feng, Hengyi, et al.
Published: (2026)
AV-RIR: Audio-Visual Room Impulse Response Estimation
by: Ratnarajah, Anton, et al.
Published: (2023)
by: Ratnarajah, Anton, et al.
Published: (2023)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
by: Choi, Jeongsoo, et al.
Published: (2023)
by: Choi, Jeongsoo, et al.
Published: (2023)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
InstructHumans: Editing Animated 3D Human Textures with Instructions
by: Zhu, Jiayin, et al.
Published: (2024)
by: Zhu, Jiayin, et al.
Published: (2024)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
ReContraster: Making Your Posters Stand Out with Regional Contrast
by: Zhang, Peixuan, et al.
Published: (2026)
by: Zhang, Peixuan, et al.
Published: (2026)
STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
Colorizing Monochromatic Radiance Fields
by: Cheng, Yean, et al.
Published: (2024)
by: Cheng, Yean, et al.
Published: (2024)
CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment
by: Hu, Yunzuo, et al.
Published: (2026)
by: Hu, Yunzuo, et al.
Published: (2026)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
by: Zhang, Jiayu, et al.
Published: (2025)
by: Zhang, Jiayu, et al.
Published: (2025)
Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
by: Chen, Sherry X., et al.
Published: (2025)
by: Chen, Sherry X., et al.
Published: (2025)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
by: Cai, Zhixi, et al.
Published: (2023)
by: Cai, Zhixi, et al.
Published: (2023)
Personalized Image Filter: Mastering Your Photographic Style
by: Zhu, Chengxuan, et al.
Published: (2025)
by: Zhu, Chengxuan, et al.
Published: (2025)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Similar Items
-
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
by: Zheng, Haojie, et al.
Published: (2025) -
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025) -
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
by: Weng, Shuchen, et al.
Published: (2024) -
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024) -
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
by: Guo, Xinyue, et al.
Published: (2025)