VALA: Learning Latent Anchors for Training-Free and Temporally Consistent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhangkai, Fan, Xuhui, Xie, Zhongyuan, Shi, Kaize, Cao, Longbing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915579662696448
author Wu, Zhangkai
Fan, Xuhui
Xie, Zhongyuan
Shi, Kaize
Cao, Longbing
author_facet Wu, Zhangkai
Fan, Xuhui
Xie, Zhongyuan
Shi, Kaize
Cao, Longbing
contents Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existing methods often rely on heuristic frame selection to maintain temporal consistency during DDIM inversion, which introduces manual bias and reduces the scalability of end-to-end inference. In this paper, we propose~\textbf{VALA} (\textbf{V}ariational \textbf{A}lignment for \textbf{L}atent \textbf{A}nchors), a variational alignment module that adaptively selects key frames and compresses their latent features into semantic anchors for consistent video editing. To learn meaningful assignments, VALA propose a variational framework with a contrastive learning objective. Therefore, it can transform cross-frame latent representations into compressed latent anchors that preserve both content and temporal coherence. Our method can be fully integrated into training-free text-to-image based video editing models. Extensive experiments on real-world video editing benchmarks show that VALA achieves state-of-the-art performance in inversion fidelity, editing quality, and temporal consistency, while offering improved efficiency over prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22970
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VALA: Learning Latent Anchors for Training-Free and Temporally Consistent
Wu, Zhangkai
Fan, Xuhui
Xie, Zhongyuan
Shi, Kaize
Cao, Longbing
Computer Vision and Pattern Recognition
Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existing methods often rely on heuristic frame selection to maintain temporal consistency during DDIM inversion, which introduces manual bias and reduces the scalability of end-to-end inference. In this paper, we propose~\textbf{VALA} (\textbf{V}ariational \textbf{A}lignment for \textbf{L}atent \textbf{A}nchors), a variational alignment module that adaptively selects key frames and compresses their latent features into semantic anchors for consistent video editing. To learn meaningful assignments, VALA propose a variational framework with a contrastive learning objective. Therefore, it can transform cross-frame latent representations into compressed latent anchors that preserve both content and temporal coherence. Our method can be fully integrated into training-free text-to-image based video editing models. Extensive experiments on real-world video editing benchmarks show that VALA achieves state-of-the-art performance in inversion fidelity, editing quality, and temporal consistency, while offering improved efficiency over prior methods.
title VALA: Learning Latent Anchors for Training-Free and Temporally Consistent
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.22970