LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Weicheng, Zhang, Zhicheng, Zhang, Zhongqi, Zhou, Juncheng, Zhu, Yongjie, Qin, Wenyu, Wang, Meng, Wan, Pengfei, Yang, Jufeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910143552159744
author Wang, Weicheng
Zhang, Zhicheng
Zhang, Zhongqi
Zhou, Juncheng
Zhu, Yongjie
Qin, Wenyu
Wang, Meng
Wan, Pengfei
Yang, Jufeng
author_facet Wang, Weicheng
Zhang, Zhicheng
Zhang, Zhongqi
Zhou, Juncheng
Zhu, Yongjie
Qin, Wenyu
Wang, Meng
Wan, Pengfei
Yang, Jufeng
contents Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data through video generation or editing models. However, compared to image editing, the high annotation costs of video data severely constrain the scale, quality, and task diversity of video editing datasets when relying on video generative models or manual annotation. To bridge this gap, we propose LIVE, a joint training framework that leverages large-scale, high-quality image editing data alongside video datasets to bolster editing capabilities. To mitigate the domain discrepancy between static images and dynamic videos, we introduce a frame-wise token noise strategy, which treats the latents of specific frames as reasoning tokens, leveraging large pretrained video generative models to create plausible temporal transformations. Moreover, through cleaning public datasets and constructing an automated data pipeline, we adopt a two-stage training strategy to anneal video editing capabilities. Furthermore, we curate a comprehensive evaluation benchmark encompassing over 60 challenging tasks that are prevalent in image editing but scarce in existing video datasets. Extensive comparative and ablation experiments demonstrate that our method achieves state-of-the-art performance. The source code will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17021
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
Wang, Weicheng
Zhang, Zhicheng
Zhang, Zhongqi
Zhou, Juncheng
Zhu, Yongjie
Qin, Wenyu
Wang, Meng
Wan, Pengfei
Yang, Jufeng
Computer Vision and Pattern Recognition
Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructing paired video editing data through video generation or editing models. However, compared to image editing, the high annotation costs of video data severely constrain the scale, quality, and task diversity of video editing datasets when relying on video generative models or manual annotation. To bridge this gap, we propose LIVE, a joint training framework that leverages large-scale, high-quality image editing data alongside video datasets to bolster editing capabilities. To mitigate the domain discrepancy between static images and dynamic videos, we introduce a frame-wise token noise strategy, which treats the latents of specific frames as reasoning tokens, leveraging large pretrained video generative models to create plausible temporal transformations. Moreover, through cleaning public datasets and constructing an automated data pipeline, we adopt a two-stage training strategy to anneal video editing capabilities. Furthermore, we curate a comprehensive evaluation benchmark encompassing over 60 challenging tasks that are prevalent in image editing but scarce in existing video datasets. Extensive comparative and ablation experiments demonstrate that our method achieves state-of-the-art performance. The source code will be publicly available.
title LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.17021