Region-Constraint In-Context Generation for Instructional Video Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhongwei, Long, Fuchen, Li, Wei, Qiu, Zhaofan, Liu, Wu, Yao, Ting, Mei, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914210502410240
author Zhang, Zhongwei
Long, Fuchen
Li, Wei
Qiu, Zhaofan
Liu, Wu
Yao, Ting
Mei, Tao
author_facet Zhang, Zhongwei
Long, Fuchen
Li, Wei
Qiu, Zhaofan
Liu, Wu
Yao, Ting
Mei, Tao
contents The In-context generation paradigm recently has demonstrated strong power in instructional image editing with both data efficiency and synthesis quality. Nevertheless, shaping such in-context learning for instruction-based video editing is not trivial. Without specifying editing regions, the results can suffer from the problem of inaccurate editing regions and the token interference between editing and non-editing areas during denoising. To address these, we present ReCo, a new instructional video editing paradigm that novelly delves into constraint modeling between editing and non-editing regions during in-context generation. Technically, ReCo width-wise concatenates source and target video for joint denoising. To calibrate video diffusion learning, ReCo capitalizes on two regularization terms, i.e., latent and attention regularization, conducting on one-step backward denoised latents and attention maps, respectively. The former increases the latent discrepancy of the editing region between source and target videos while reducing that of non-editing areas, emphasizing the modification on editing area and alleviating outside unexpected content generation. The latter suppresses the attention of tokens in the editing region to the tokens in counterpart of the source video, thereby mitigating their interference during novel object generation in target video. Furthermore, we propose a large-scale, high-quality video editing dataset, i.e., ReCo-Data, comprising 500K instruction-video pairs to benefit model training. Extensive experiments conducted on four major instruction-based video editing tasks demonstrate the superiority of our proposal.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Region-Constraint In-Context Generation for Instructional Video Editing
Zhang, Zhongwei
Long, Fuchen
Li, Wei
Qiu, Zhaofan
Liu, Wu
Yao, Ting
Mei, Tao
Computer Vision and Pattern Recognition
Multimedia
The In-context generation paradigm recently has demonstrated strong power in instructional image editing with both data efficiency and synthesis quality. Nevertheless, shaping such in-context learning for instruction-based video editing is not trivial. Without specifying editing regions, the results can suffer from the problem of inaccurate editing regions and the token interference between editing and non-editing areas during denoising. To address these, we present ReCo, a new instructional video editing paradigm that novelly delves into constraint modeling between editing and non-editing regions during in-context generation. Technically, ReCo width-wise concatenates source and target video for joint denoising. To calibrate video diffusion learning, ReCo capitalizes on two regularization terms, i.e., latent and attention regularization, conducting on one-step backward denoised latents and attention maps, respectively. The former increases the latent discrepancy of the editing region between source and target videos while reducing that of non-editing areas, emphasizing the modification on editing area and alleviating outside unexpected content generation. The latter suppresses the attention of tokens in the editing region to the tokens in counterpart of the source video, thereby mitigating their interference during novel object generation in target video. Furthermore, we propose a large-scale, high-quality video editing dataset, i.e., ReCo-Data, comprising 500K instruction-video pairs to benefit model training. Extensive experiments conducted on four major instruction-based video editing tasks demonstrate the superiority of our proposal.
title Region-Constraint In-Context Generation for Instructional Video Editing
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2512.17650