AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Haojie, Weng, Shuchen, Liu, Jingqi, Yang, Siqi, Shi, Boxin, Wang, Xinlong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
by: Zheng, Haojie, et al.
Published: (2026)
by: Zheng, Haojie, et al.
Published: (2026)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
by: Weng, Shuchen, et al.
Published: (2024)
by: Weng, Shuchen, et al.
Published: (2024)
PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
L-C4: Language-Based Video Colorization for Creative and Consistent Color
by: Chang, Zheng, et al.
Published: (2024)
by: Chang, Zheng, et al.
Published: (2024)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
by: Cai, Ziqi, et al.
Published: (2026)
by: Cai, Ziqi, et al.
Published: (2026)
Affective Image Editing: Shaping Emotional Factors via Text Descriptions
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
by: Zhang, Peixuan, et al.
Published: (2026)
by: Zhang, Peixuan, et al.
Published: (2026)
AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance
by: Zhou, Xilong, et al.
Published: (2026)
by: Zhou, Xilong, et al.
Published: (2026)
A$^2$-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
by: Zheng, Huayu, et al.
Published: (2026)
by: Zheng, Huayu, et al.
Published: (2026)
Language-guided Image Reflection Separation
by: Zhong, Haofeng, et al.
Published: (2024)
by: Zhong, Haofeng, et al.
Published: (2024)
Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
by: Zhang, Peixuan, et al.
Published: (2025)
by: Zhang, Peixuan, et al.
Published: (2025)
Edit as You See: Image-guided Video Editing via Masked Motion Modeling
by: Huang, Zhi-Lin, et al.
Published: (2025)
by: Huang, Zhi-Lin, et al.
Published: (2025)
MDE-Edit: Masked Dual-Editing for Multi-Object Image Editing via Diffusion Models
by: Zhu, Hongyang, et al.
Published: (2025)
by: Zhu, Hongyang, et al.
Published: (2025)
FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing
by: Koo, Gwanhyeong, et al.
Published: (2024)
by: Koo, Gwanhyeong, et al.
Published: (2024)
ReContraster: Making Your Posters Stand Out with Regional Contrast
by: Zhang, Peixuan, et al.
Published: (2026)
by: Zhang, Peixuan, et al.
Published: (2026)
Colorizing Monochromatic Radiance Fields
by: Cheng, Yean, et al.
Published: (2024)
by: Cheng, Yean, et al.
Published: (2024)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
by: Xu, Zitong, et al.
Published: (2026)
by: Xu, Zitong, et al.
Published: (2026)
LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
by: Gao, Chenjian, et al.
Published: (2025)
by: Gao, Chenjian, et al.
Published: (2025)
Personalized Image Filter: Mastering Your Photographic Style
by: Zhu, Chengxuan, et al.
Published: (2025)
by: Zhu, Chengxuan, et al.
Published: (2025)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026)
by: Somayazulu, Arjun, et al.
Published: (2026)
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
by: Couairon, Paul, et al.
Published: (2023)
by: Couairon, Paul, et al.
Published: (2023)
Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
by: Cai, Honghao, et al.
Published: (2026)
by: Cai, Honghao, et al.
Published: (2026)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
by: Sun, Qianqian, et al.
Published: (2025)
by: Sun, Qianqian, et al.
Published: (2025)
Semantic Granularity Navigation in Image Editing
by: Lu, Liangsi, et al.
Published: (2026)
by: Lu, Liangsi, et al.
Published: (2026)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing
by: Yuan, Tianshuo, et al.
Published: (2024)
by: Yuan, Tianshuo, et al.
Published: (2024)
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
by: Zuo, Yi, et al.
Published: (2024)
by: Zuo, Yi, et al.
Published: (2024)
Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
by: Shi, Yiqing, et al.
Published: (2025)
by: Shi, Yiqing, et al.
Published: (2025)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
Group Editing: Edit Multiple Images in One Go
by: Ma, Yue, et al.
Published: (2026)
by: Ma, Yue, et al.
Published: (2026)
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
by: Litman, Yehonathan, et al.
Published: (2026)
by: Litman, Yehonathan, et al.
Published: (2026)
Granularity-Aware Transfer for Tree Instance Segmentation in Synthetic and Real Forests
by: Deoli, Pankaj, et al.
Published: (2026)
by: Deoli, Pankaj, et al.
Published: (2026)
Data standardization for robust lip sync
by: Wang, Chun
Published: (2022)
by: Wang, Chun
Published: (2022)
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
by: Zhang, Youyuan, et al.
Published: (2024)
by: Zhang, Youyuan, et al.
Published: (2024)
iMOVE: Instance-Motion-Aware Video Understanding
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
Similar Items
-
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
by: Zheng, Haojie, et al.
Published: (2026) -
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025) -
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
by: Weng, Shuchen, et al.
Published: (2024) -
PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
by: Xia, Yifei, et al.
Published: (2025) -
L-C4: Language-Based Video Colorization for Creative and Consistent Color
by: Chang, Zheng, et al.
Published: (2024)