Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Yizhou, Liu, Zihua, Monno, Yusuke, Okutomi, Masatoshi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2501.02269
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916551026802688
author Li, Yizhou
Liu, Zihua
Monno, Yusuke
Okutomi, Masatoshi
author_facet Li, Yizhou
Liu, Zihua
Monno, Yusuke
Okutomi, Masatoshi
contents In this paper, we propose the first diffusion-based all-in-one video restoration method that utilizes the power of a pre-trained Stable Diffusion and a fine-tuned ControlNet. Our method can restore various types of video degradation with a single unified model, overcoming the limitation of standard methods that require specific models for each restoration task. Our contributions include an efficient training strategy with Task Prompt Guidance (TPG) for diverse restoration tasks, an inference strategy that combines Denoising Diffusion Implicit Models~(DDIM) inversion with a novel Sliding Window Cross-Frame Attention (SW-CFA) mechanism for enhanced content preservation and temporal consistency, and a scalable pipeline that makes our method all-in-one to adapt to different video restoration tasks. Through extensive experiments on five video restoration tasks, we demonstrate the superiority of our method in generalization capability to real-world videos and temporal consistency preservation over existing state-of-the-art methods. Our method advances the video restoration task by providing a unified solution that enhances video quality across multiple applications.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration
Li, Yizhou
Liu, Zihua
Monno, Yusuke
Okutomi, Masatoshi
Computer Vision and Pattern Recognition
In this paper, we propose the first diffusion-based all-in-one video restoration method that utilizes the power of a pre-trained Stable Diffusion and a fine-tuned ControlNet. Our method can restore various types of video degradation with a single unified model, overcoming the limitation of standard methods that require specific models for each restoration task. Our contributions include an efficient training strategy with Task Prompt Guidance (TPG) for diverse restoration tasks, an inference strategy that combines Denoising Diffusion Implicit Models~(DDIM) inversion with a novel Sliding Window Cross-Frame Attention (SW-CFA) mechanism for enhanced content preservation and temporal consistency, and a scalable pipeline that makes our method all-in-one to adapt to different video restoration tasks. Through extensive experiments on five video restoration tasks, we demonstrate the superiority of our method in generalization capability to real-world videos and temporal consistency preservation over existing state-of-the-art methods. Our method advances the video restoration task by providing a unified solution that enhances video quality across multiple applications.
title TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.02269