InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yuhui, Chen, Liyi, Li, Ruibin, Wang, Shihao, Xie, Chenxi, Zhang, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913938217631744
author Wu, Yuhui
Chen, Liyi
Li, Ruibin
Wang, Shihao
Xie, Chenxi
Zhang, Lei
author_facet Wu, Yuhui
Chen, Liyi
Li, Ruibin
Wang, Shihao
Xie, Chenxi
Zhang, Lei
contents Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video, instruction) is a challenging task. Existing datasets mostly consist of low-resolution, short duration, and limited amount of source videos with unsatisfactory editing quality, limiting the performance of trained editing models. In this work, we present a high-quality Instruction-based Video Editing dataset with 1M triplets, namely InsViE-1M. We first curate high-resolution and high-quality source videos and images, then design an effective editing-filtering pipeline to construct high-quality editing triplets for model training. For a source video, we generate multiple edited samples of its first frame with different intensities of classifier-free guidance, which are automatically filtered by GPT-4o with carefully crafted guidelines. The edited first frame is propagated to subsequent frames to produce the edited video, followed by another round of filtering for frame quality and motion evaluation. We also generate and filter a variety of video editing triplets from high-quality images. With the InsViE-1M dataset, we propose a multi-stage learning strategy to train our InsViE model, progressively enhancing its instruction following and editing ability. Extensive experiments demonstrate the advantages of our InsViE-1M dataset and the trained model over state-of-the-art works. Codes are available at \href{https://github.com/langmanbusi/InsViE}{InsViE}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20287
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
Wu, Yuhui
Chen, Liyi
Li, Ruibin
Wang, Shihao
Xie, Chenxi
Zhang, Lei
Computer Vision and Pattern Recognition
Instruction-based video editing allows effective and interactive editing of videos using only instructions without extra inputs such as masks or attributes. However, collecting high-quality training triplets (source video, edited video, instruction) is a challenging task. Existing datasets mostly consist of low-resolution, short duration, and limited amount of source videos with unsatisfactory editing quality, limiting the performance of trained editing models. In this work, we present a high-quality Instruction-based Video Editing dataset with 1M triplets, namely InsViE-1M. We first curate high-resolution and high-quality source videos and images, then design an effective editing-filtering pipeline to construct high-quality editing triplets for model training. For a source video, we generate multiple edited samples of its first frame with different intensities of classifier-free guidance, which are automatically filtered by GPT-4o with carefully crafted guidelines. The edited first frame is propagated to subsequent frames to produce the edited video, followed by another round of filtering for frame quality and motion evaluation. We also generate and filter a variety of video editing triplets from high-quality images. With the InsViE-1M dataset, we propose a multi-stage learning strategy to train our InsViE model, progressively enhancing its instruction following and editing ability. Extensive experiments demonstrate the advantages of our InsViE-1M dataset and the trained model over state-of-the-art works. Codes are available at \href{https://github.com/langmanbusi/InsViE}{InsViE}.
title InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.20287