InstructVEdit: A Holistic Approach for Instructional Video Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Chi, Feng, Chengjian, Yan, Feng, Zhang, Qiming, Zhang, Mingjin, Zhong, Yujie, Zhang, Jing, Ma, Lin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912287090016256
author Zhang, Chi
Feng, Chengjian
Yan, Feng
Zhang, Qiming
Zhang, Mingjin
Zhong, Yujie
Zhang, Jing
Ma, Lin
author_facet Zhang, Chi
Feng, Chengjian
Yan, Feng
Zhang, Qiming
Zhang, Mingjin
Zhong, Yujie
Zhang, Jing
Ma, Lin
contents Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the systematic exploration of model architectures and training strategies. While prior work has improved specific aspects of video editing (e.g., synthesizing a video dataset using image editing techniques or decomposed video editing training), a holistic framework addressing the above challenges remains underexplored. In this study, we introduce InstructVEdit, a full-cycle instructional video editing approach that: (1) establishes a reliable dataset curation workflow to initialize training, (2) incorporates two model architectural improvements to enhance edit quality while preserving temporal consistency, and (3) proposes an iterative refinement strategy leveraging real-world data to enhance generalization and minimize train-test discrepancies. Extensive experiments show that InstructVEdit achieves state-of-the-art performance in instruction-based video editing, demonstrating robust adaptability to diverse real-world scenarios. Project page: https://o937-blip.github.io/InstructVEdit.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17641
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InstructVEdit: A Holistic Approach for Instructional Video Editing
Zhang, Chi
Feng, Chengjian
Yan, Feng
Zhang, Qiming
Zhang, Mingjin
Zhong, Yujie
Zhang, Jing
Ma, Lin
Computer Vision and Pattern Recognition
Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the systematic exploration of model architectures and training strategies. While prior work has improved specific aspects of video editing (e.g., synthesizing a video dataset using image editing techniques or decomposed video editing training), a holistic framework addressing the above challenges remains underexplored. In this study, we introduce InstructVEdit, a full-cycle instructional video editing approach that: (1) establishes a reliable dataset curation workflow to initialize training, (2) incorporates two model architectural improvements to enhance edit quality while preserving temporal consistency, and (3) proposes an iterative refinement strategy leveraging real-world data to enhance generalization and minimize train-test discrepancies. Extensive experiments show that InstructVEdit achieves state-of-the-art performance in instruction-based video editing, demonstrating robust adaptability to diverse real-world scenarios. Project page: https://o937-blip.github.io/InstructVEdit.
title InstructVEdit: A Holistic Approach for Instructional Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.17641