Low-Cost Test-Time Adaptation for Robust Video Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jianhui, Chen, Yinda, He, Yangfan, Song, Xinyuan, Xin, Yi, Zhang, Dapeng, Wan, Zhongwei, Li, Bin, Zhang, Rongchao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915416940478464
author Wang, Jianhui
Chen, Yinda
He, Yangfan
Song, Xinyuan
Xin, Yi
Zhang, Dapeng
Wan, Zhongwei
Li, Bin
Zhang, Rongchao
author_facet Wang, Jianhui
Chen, Yinda
He, Yangfan
Song, Xinyuan
Xin, Yi
Zhang, Dapeng
Wan, Zhongwei
Li, Bin
Zhang, Rongchao
contents Video editing is a critical component of content creation that transforms raw footage into coherent works aligned with specific visual and narrative objectives. Existing approaches face two major challenges: temporal inconsistencies due to failure in capturing complex motion patterns, and overfitting to simple prompts arising from limitations in UNet backbone architectures. While learning-based methods can enhance editing quality, they typically demand substantial computational resources and are constrained by the scarcity of high-quality annotated data. In this paper, we present Vid-TTA, a lightweight test-time adaptation framework that personalizes optimization for each test video during inference through self-supervised auxiliary tasks. Our approach incorporates a motion-aware frame reconstruction mechanism that identifies and preserves crucial movement regions, alongside a prompt perturbation and reconstruction strategy that strengthens model robustness to diverse textual descriptions. These innovations are orchestrated by a meta-learning driven dynamic loss balancing mechanism that adaptively adjusts the optimization process based on video characteristics. Extensive experiments demonstrate that Vid-TTA significantly improves video temporal consistency and mitigates prompt overfitting while maintaining low computational overhead, offering a plug-and-play performance boost for existing video editing models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21858
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Low-Cost Test-Time Adaptation for Robust Video Editing
Wang, Jianhui
Chen, Yinda
He, Yangfan
Song, Xinyuan
Xin, Yi
Zhang, Dapeng
Wan, Zhongwei
Li, Bin
Zhang, Rongchao
Computer Vision and Pattern Recognition
Video editing is a critical component of content creation that transforms raw footage into coherent works aligned with specific visual and narrative objectives. Existing approaches face two major challenges: temporal inconsistencies due to failure in capturing complex motion patterns, and overfitting to simple prompts arising from limitations in UNet backbone architectures. While learning-based methods can enhance editing quality, they typically demand substantial computational resources and are constrained by the scarcity of high-quality annotated data. In this paper, we present Vid-TTA, a lightweight test-time adaptation framework that personalizes optimization for each test video during inference through self-supervised auxiliary tasks. Our approach incorporates a motion-aware frame reconstruction mechanism that identifies and preserves crucial movement regions, alongside a prompt perturbation and reconstruction strategy that strengthens model robustness to diverse textual descriptions. These innovations are orchestrated by a meta-learning driven dynamic loss balancing mechanism that adaptively adjusts the optimization process based on video characteristics. Extensive experiments demonstrate that Vid-TTA significantly improves video temporal consistency and mitigates prompt overfitting while maintaining low computational overhead, offering a plug-and-play performance boost for existing video editing models.
title Low-Cost Test-Time Adaptation for Robust Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.21858