Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Bohai, Luo, Hao, Guo, Song, Dong, Peiran, Zhou, Qihua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909533376348160
author Gu, Bohai
Luo, Hao
Guo, Song
Dong, Peiran
Zhou, Qihua
author_facet Gu, Bohai
Luo, Hao
Guo, Song
Dong, Peiran
Zhou, Qihua
contents The text-guided video inpainting technique has significantly improved the performance of content generation applications. A recent family for these improvements uses diffusion models, which have become essential for achieving high-quality video inpainting results, yet they still face performance bottlenecks in temporal consistency and computational efficiency. This motivates us to propose a new video inpainting framework using optical Flow-guided Efficient Diffusion (FloED) for higher video coherence. Specifically, FloED employs a dual-branch architecture, where the time-agnostic flow branch restores corrupted flow first, and the multi-scale flow adapters provide motion guidance to the main inpainting branch. Besides, a training-free latent interpolation method is proposed to accelerate the multi-step denoising process using flow warping. With the flow attention cache mechanism, FLoED efficiently reduces the computational cost of incorporating optical flow. Extensive experiments on background restoration and object removal tasks show that FloED outperforms state-of-the-art diffusion-based methods in both quality and efficiency. Our codes and models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00857
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
Gu, Bohai
Luo, Hao
Guo, Song
Dong, Peiran
Zhou, Qihua
Computer Vision and Pattern Recognition
The text-guided video inpainting technique has significantly improved the performance of content generation applications. A recent family for these improvements uses diffusion models, which have become essential for achieving high-quality video inpainting results, yet they still face performance bottlenecks in temporal consistency and computational efficiency. This motivates us to propose a new video inpainting framework using optical Flow-guided Efficient Diffusion (FloED) for higher video coherence. Specifically, FloED employs a dual-branch architecture, where the time-agnostic flow branch restores corrupted flow first, and the multi-scale flow adapters provide motion guidance to the main inpainting branch. Besides, a training-free latent interpolation method is proposed to accelerate the multi-step denoising process using flow warping. With the flow attention cache mechanism, FLoED efficiently reduces the computational cost of incorporating optical flow. Extensive experiments on background restoration and object removal tasks show that FloED outperforms state-of-the-art diffusion-based methods in both quality and efficiency. Our codes and models will be made publicly available.
title Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00857