Temporal-Consistent Video Restoration with Pre-trained Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hengkang, Liu, Yang, Liu, Huidong, Wang, Chien-Chih, Guo, Yanhui, Li, Hongdong, Wang, Bryan, Sun, Ju
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912282444824576
author Wang, Hengkang
Liu, Yang
Liu, Huidong
Wang, Chien-Chih
Guo, Yanhui
Li, Hongdong
Wang, Bryan
Sun, Ju
author_facet Wang, Hengkang
Liu, Yang
Liu, Huidong
Wang, Chien-Chih
Guo, Yanhui
Li, Hongdong
Wang, Bryan
Sun, Ju
contents Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with 3D video data, VR is inherently computationally intensive. In this paper, we advocate viewing the reverse process in DMs as a function and present a novel Maximum a Posterior (MAP) framework that directly parameterizes video frames in the seed space of DMs, eliminating approximation errors. We also introduce strategies to promote bilevel temporal consistency: semantic consistency by leveraging clustering structures in the seed space, and pixel-level consistency by progressive warping with optical flow refinements. Extensive experiments on multiple virtual reality tasks demonstrate superior visual quality and temporal consistency achieved by our method compared to the state-of-the-art.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
Wang, Hengkang
Liu, Yang
Liu, Huidong
Wang, Chien-Chih
Guo, Yanhui
Li, Hongdong
Wang, Bryan
Sun, Ju
Computer Vision and Pattern Recognition
Video restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with 3D video data, VR is inherently computationally intensive. In this paper, we advocate viewing the reverse process in DMs as a function and present a novel Maximum a Posterior (MAP) framework that directly parameterizes video frames in the seed space of DMs, eliminating approximation errors. We also introduce strategies to promote bilevel temporal consistency: semantic consistency by leveraging clustering structures in the seed space, and pixel-level consistency by progressive warping with optical flow refinements. Extensive experiments on multiple virtual reality tasks demonstrate superior visual quality and temporal consistency achieved by our method compared to the state-of-the-art.
title Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.14863