EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Litman, Yehonathan, Liu, Shikun, Seyb, Dario, Milef, Nicholas, Zhou, Yang, Marshall, Carl, Tulsiani, Shubham, Leak, Caleb
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915903940067328
author Litman, Yehonathan
Liu, Shikun
Seyb, Dario
Milef, Nicholas
Zhou, Yang
Marshall, Carl
Tulsiani, Shubham
Leak, Caleb
author_facet Litman, Yehonathan
Liu, Shikun
Seyb, Dario
Milef, Nicholas
Zhou, Yang
Marshall, Carl
Tulsiani, Shubham
Leak, Caleb
contents High-fidelity generative video editing has seen significant quality improvements by leveraging pre-trained video foundation models. However, their computational cost is a major bottleneck, as they are often designed to inefficiently process the full video context regardless of the inpainting mask's size, even for sparse, localized edits. In this paper, we introduce EditCtrl, an efficient video inpainting control framework that focuses computation only where it is needed. Our approach features a novel local video context module that operates solely on masked tokens, yielding a computational cost proportional to the edit size. This local-first generation is then guided by a lightweight temporal global context embedder that ensures video-wide context consistency with minimal overhead. Not only is EditCtrl 10 times more compute efficient than state-of-the-art generative editing methods, it even improves editing quality compared to methods designed with full-attention. Finally, we showcase how EditCtrl unlocks new capabilities, including multi-region editing with text prompts and autoregressive content propagation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15031
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
Litman, Yehonathan
Liu, Shikun
Seyb, Dario
Milef, Nicholas
Zhou, Yang
Marshall, Carl
Tulsiani, Shubham
Leak, Caleb
Computer Vision and Pattern Recognition
High-fidelity generative video editing has seen significant quality improvements by leveraging pre-trained video foundation models. However, their computational cost is a major bottleneck, as they are often designed to inefficiently process the full video context regardless of the inpainting mask's size, even for sparse, localized edits. In this paper, we introduce EditCtrl, an efficient video inpainting control framework that focuses computation only where it is needed. Our approach features a novel local video context module that operates solely on masked tokens, yielding a computational cost proportional to the edit size. This local-first generation is then guided by a lightweight temporal global context embedder that ensures video-wide context consistency with minimal overhead. Not only is EditCtrl 10 times more compute efficient than state-of-the-art generative editing methods, it even improves editing quality compared to methods designed with full-attention. Finally, we showcase how EditCtrl unlocks new capabilities, including multi-region editing with text prompts and autoregressive content propagation.
title EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.15031