Versatile Editing of Video Content, Actions, and Dynamics without Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kulikov, Vladimir, Paiss, Roni, Voynov, Andrey, Mosseri, Inbar, Dekel, Tali, Michaeli, Tomer
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917351873576960
author Kulikov, Vladimir
Paiss, Roni
Voynov, Andrey
Mosseri, Inbar
Dekel, Tali
Michaeli, Tomer
author_facet Kulikov, Vladimir
Paiss, Roni
Voynov, Andrey
Mosseri, Inbar
Dekel, Tali
Michaeli, Tomer
contents Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of other objects in real-world videos, remains a major challenge. Existing trained models struggle with complex edits, likely due to the difficulty of collecting relevant training data. Similarly, existing training-free methods are inherently restricted to structure- and motion-preserving edits and do not support modification of motion or interactions. Here, we introduce DynaEdit, a training-free editing method that unlocks versatile video editing capabilities with pretrained text-to-video flow models. Our method relies on the recently introduced inversion-free approach, which does not intervene in the model internals, and is thus model-agnostic. We show that naively attempting to adapt this approach to general unconstrained editing results in severe low-frequency misalignment and high-frequency jitter. We explain the sources for these phenomena and introduce novel mechanisms for overcoming them. Through extensive experiments, we show that DynaEdit achieves state-of-the-art results on complex text-based video editing tasks, including modifying actions, inserting objects that interact with the scene, and introducing global effects.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17989
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Versatile Editing of Video Content, Actions, and Dynamics without Training
Kulikov, Vladimir
Paiss, Roni
Voynov, Andrey
Mosseri, Inbar
Dekel, Tali
Michaeli, Tomer
Computer Vision and Pattern Recognition
Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of other objects in real-world videos, remains a major challenge. Existing trained models struggle with complex edits, likely due to the difficulty of collecting relevant training data. Similarly, existing training-free methods are inherently restricted to structure- and motion-preserving edits and do not support modification of motion or interactions. Here, we introduce DynaEdit, a training-free editing method that unlocks versatile video editing capabilities with pretrained text-to-video flow models. Our method relies on the recently introduced inversion-free approach, which does not intervene in the model internals, and is thus model-agnostic. We show that naively attempting to adapt this approach to general unconstrained editing results in severe low-frequency misalignment and high-frequency jitter. We explain the sources for these phenomena and introduce novel mechanisms for overcoming them. Through extensive experiments, we show that DynaEdit achieves state-of-the-art results on complex text-based video editing tasks, including modifying actions, inserting objects that interact with the scene, and introducing global effects.
title Versatile Editing of Video Content, Actions, and Dynamics without Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.17989