Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Cong, Yue, Huanjing, Xie, Shangbin, Liu, Xin, Yang, Jingyu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912859760361472
author Cao, Cong
Yue, Huanjing
Xie, Shangbin
Liu, Xin
Yang, Jingyu
author_facet Cao, Cong
Yue, Huanjing
Xie, Shangbin
Liu, Xin
Yang, Jingyu
contents Although diffusion-based zero-shot image restoration and enhancement methods have achieved great success, applying them to video restoration or enhancement will lead to severe temporal flickering. In this paper, we propose the first framework that utilizes the rapidly-developed video diffusion model to assist the image-based method in maintaining more temporal consistency for zero-shot video restoration and enhancement. We propose homologous latents fusion, heterogenous latents fusion, and a COT-based fusion ratio strategy to utilize both homologous and heterogenous text-to-video diffusion models to complement the image method. Moreover, we propose temporal-strengthening post-processing to utilize the image-to-video diffusion model to further improve temporal consistency. Our method is training-free and can be applied to any diffusion-based image restoration and enhancement methods. Experimental results demonstrate the superiority of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21922
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
Cao, Cong
Yue, Huanjing
Xie, Shangbin
Liu, Xin
Yang, Jingyu
Computer Vision and Pattern Recognition
Although diffusion-based zero-shot image restoration and enhancement methods have achieved great success, applying them to video restoration or enhancement will lead to severe temporal flickering. In this paper, we propose the first framework that utilizes the rapidly-developed video diffusion model to assist the image-based method in maintaining more temporal consistency for zero-shot video restoration and enhancement. We propose homologous latents fusion, heterogenous latents fusion, and a COT-based fusion ratio strategy to utilize both homologous and heterogenous text-to-video diffusion models to complement the image method. Moreover, we propose temporal-strengthening post-processing to utilize the image-to-video diffusion model to further improve temporal consistency. Our method is training-free and can be applied to any diffusion-based image restoration and enhancement methods. Experimental results demonstrate the superiority of the proposed method.
title Zero-Shot Video Restoration and Enhancement with Assistance of Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.21922