Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Qingyan, Wang, Qiuyu, Ouyang, Hao, Yu, Yue, Wang, Hanlin, Wang, Wen, Cheng, Ka Leong, Ma, Shuailei, Zeng, Yanhong, Liu, Zichen, Xu, Yinghao, Shen, Yujun, Chen, Qifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918251544444928
author Bai, Qingyan
Wang, Qiuyu
Ouyang, Hao
Yu, Yue
Wang, Hanlin
Wang, Wen
Cheng, Ka Leong
Ma, Shuailei
Zeng, Yanhong
Liu, Zichen
Xu, Yinghao
Shen, Yujun
Chen, Qifeng
author_facet Bai, Qingyan
Wang, Qiuyu
Ouyang, Hao
Yu, Yue
Wang, Hanlin
Wang, Wen
Cheng, Ka Leong
Ma, Shuailei
Zeng, Yanhong
Liu, Zichen
Xu, Yinghao
Shen, Yujun
Chen, Qifeng
contents Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence. Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy. The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15742
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
Bai, Qingyan
Wang, Qiuyu
Ouyang, Hao
Yu, Yue
Wang, Hanlin
Wang, Wen
Cheng, Ka Leong
Ma, Shuailei
Zeng, Yanhong
Liu, Zichen
Xu, Yinghao
Shen, Yujun
Chen, Qifeng
Computer Vision and Pattern Recognition
Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence. Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy. The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing.
title Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.15742