Enhancing Video Inpainting with Aligned Frame Interval Guidance

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Ming, Yu, Junqiu, Dong, Qiaole, Xue, Xiangyang, Fu, Yanwei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918201925828608
author Xie, Ming
Yu, Junqiu
Dong, Qiaole
Xue, Xiangyang
Fu, Yanwei
author_facet Xie, Ming
Yu, Junqiu
Dong, Qiaole
Xue, Xiangyang
Fu, Yanwei
contents Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. Nevertheless, these methods suffer from severe content degradation within video chunks. Furthermore, the absence of a robust frame alignment scheme compromises intra-chunk and inter-chunk spatiotemporal stability, resulting in insufficient control over the entire video. To address these limitations, we propose VidPivot, a novel framework that decouples video inpainting into two sub-tasks: multi-frame consistent image inpainting and masked area motion propagation. Our approach introduces frame interval priors as spatiotemporal cues to guide the inpainting process. To enhance cross-frame coherence, we design a FrameProp Module that implements a frame content propagation strategy, diffusing reference frame content into subsequent frames via a splicing mechanism. Additionally, a dedicated context controller encodes these coherent frame priors into the I2V generative backbone, effectively serving as soft constrain to suppress content distortion during generation. Extensive evaluations demonstrate that VidPivot achieves competitive performance across diverse benchmarks and generalizes well to different video inpainting scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21461
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Video Inpainting with Aligned Frame Interval Guidance
Xie, Ming
Yu, Junqiu
Dong, Qiaole
Xue, Xiangyang
Fu, Yanwei
Computer Vision and Pattern Recognition
Recent image-to-video (I2V) based video inpainting methods have made significant strides by leveraging single-image priors and modeling temporal consistency across masked frames. Nevertheless, these methods suffer from severe content degradation within video chunks. Furthermore, the absence of a robust frame alignment scheme compromises intra-chunk and inter-chunk spatiotemporal stability, resulting in insufficient control over the entire video. To address these limitations, we propose VidPivot, a novel framework that decouples video inpainting into two sub-tasks: multi-frame consistent image inpainting and masked area motion propagation. Our approach introduces frame interval priors as spatiotemporal cues to guide the inpainting process. To enhance cross-frame coherence, we design a FrameProp Module that implements a frame content propagation strategy, diffusing reference frame content into subsequent frames via a splicing mechanism. Additionally, a dedicated context controller encodes these coherent frame priors into the I2V generative backbone, effectively serving as soft constrain to suppress content distortion during generation. Extensive evaluations demonstrate that VidPivot achieves competitive performance across diverse benchmarks and generalizes well to different video inpainting scenarios.
title Enhancing Video Inpainting with Aligned Frame Interval Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.21461