Saved in:
Bibliographic Details
Main Authors: He, Haoyang, Wang, Jie, Zhang, Jiangning, Xue, Zhucun, Bu, Xingyuan, Yang, Qiangpeng, Wen, Shilei, Xie, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.07826
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917148168814592
author He, Haoyang
Wang, Jie
Zhang, Jiangning
Xue, Zhucun
Bu, Xingyuan
Yang, Qiangpeng
Wen, Shilei
Xie, Lei
author_facet He, Haoyang
Wang, Jie
Zhang, Jiangning
Xue, Zhucun
Bu, Xingyuan
Yang, Qiangpeng
Wen, Shilei
Xie, Lei
contents The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an open-source, large-scale, and high-quality dataset for instruction-based video editing. It comprises two primary categories: spatially-aligned edits (Global Style, Background Change, Local Change, Local Remove, Local Add, and Subtitles Edit) and non-spatially-aligned edits (Camera Multi-Shot Edit and Creative Edit). All edit types are generated via a meticulously designed data pipeline with rigorous quality filtering. OpenVE-3M surpasses existing open-source datasets in terms of scale, diversity of edit types, instruction length, and overall quality. Furthermore, to address the lack of a unified benchmark in the field, we construct OpenVE-Bench, containing 431 video-edit pairs that cover a diverse range of editing tasks with three key metrics highly aligned with human judgment. We present OpenVE-Edit, a 5B model trained on our dataset that demonstrates remarkable efficiency and effectiveness by setting a new state-of-the-art on OpenVE-Bench, outperforming all prior open-source models including a 14B baseline. Project page is at https://lewandofskee.github.io/projects/OpenVE.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07826
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
He, Haoyang
Wang, Jie
Zhang, Jiangning
Xue, Zhucun
Bu, Xingyuan
Yang, Qiangpeng
Wen, Shilei
Xie, Lei
Computer Vision and Pattern Recognition
The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an open-source, large-scale, and high-quality dataset for instruction-based video editing. It comprises two primary categories: spatially-aligned edits (Global Style, Background Change, Local Change, Local Remove, Local Add, and Subtitles Edit) and non-spatially-aligned edits (Camera Multi-Shot Edit and Creative Edit). All edit types are generated via a meticulously designed data pipeline with rigorous quality filtering. OpenVE-3M surpasses existing open-source datasets in terms of scale, diversity of edit types, instruction length, and overall quality. Furthermore, to address the lack of a unified benchmark in the field, we construct OpenVE-Bench, containing 431 video-edit pairs that cover a diverse range of editing tasks with three key metrics highly aligned with human judgment. We present OpenVE-Edit, a 5B model trained on our dataset that demonstrates remarkable efficiency and effectiveness by setting a new state-of-the-art on OpenVE-Bench, outperforming all prior open-source models including a 14B baseline. Project page is at https://lewandofskee.github.io/projects/OpenVE.
title OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.07826