Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rao, Chen, Li, Guangyuan, Lan, Zehua, Sun, Jiakai, Luan, Junsheng, Xing, Wei, Zhao, Lei, Lin, Huaizhong, Dong, Jianfeng, Zhang, Dalong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914922382753792
author Rao, Chen
Li, Guangyuan
Lan, Zehua
Sun, Jiakai
Luan, Junsheng
Xing, Wei
Zhao, Lei
Lin, Huaizhong
Dong, Jianfeng
Zhang, Dalong
author_facet Rao, Chen
Li, Guangyuan
Lan, Zehua
Sun, Jiakai
Luan, Junsheng
Xing, Wei
Zhao, Lei
Lin, Huaizhong
Dong, Jianfeng
Zhang, Dalong
contents Current video deblurring methods have limitations in recovering high-frequency information since the regression losses are conservative with high-frequency details. Since Diffusion Models (DMs) have strong capabilities in generating high-frequency details, we consider introducing DMs into the video deblurring task. However, we found that directly applying DMs to the video deblurring task has the following problems: (1) DMs require many iteration steps to generate videos from Gaussian noise, which consumes many computational resources. (2) DMs are easily misled by the blurry artifacts in the video, resulting in irrational content and distortion of the deblurred video. To address the above issues, we propose a novel video deblurring framework VD-Diff that integrates the diffusion model into the Wavelet-Aware Dynamic Transformer (WADT). Specifically, we perform the diffusion model in a highly compact latent space to generate prior features containing high-frequency information that conforms to the ground truth distribution. We design the WADT to preserve and recover the low-frequency information in the video while utilizing the high-frequency information generated by the diffusion model. Extensive experiments show that our proposed VD-Diff outperforms SOTA methods on GoPro, DVD, BSD, and Real-World Video datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13459
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model
Rao, Chen
Li, Guangyuan
Lan, Zehua
Sun, Jiakai
Luan, Junsheng
Xing, Wei
Zhao, Lei
Lin, Huaizhong
Dong, Jianfeng
Zhang, Dalong
Computer Vision and Pattern Recognition
I.4.4
Current video deblurring methods have limitations in recovering high-frequency information since the regression losses are conservative with high-frequency details. Since Diffusion Models (DMs) have strong capabilities in generating high-frequency details, we consider introducing DMs into the video deblurring task. However, we found that directly applying DMs to the video deblurring task has the following problems: (1) DMs require many iteration steps to generate videos from Gaussian noise, which consumes many computational resources. (2) DMs are easily misled by the blurry artifacts in the video, resulting in irrational content and distortion of the deblurred video. To address the above issues, we propose a novel video deblurring framework VD-Diff that integrates the diffusion model into the Wavelet-Aware Dynamic Transformer (WADT). Specifically, we perform the diffusion model in a highly compact latent space to generate prior features containing high-frequency information that conforms to the ground truth distribution. We design the WADT to preserve and recover the low-frequency information in the video while utilizing the high-frequency information generated by the diffusion model. Extensive experiments show that our proposed VD-Diff outperforms SOTA methods on GoPro, DVD, BSD, and Real-World Video datasets.
title Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model
topic Computer Vision and Pattern Recognition
I.4.4
url https://arxiv.org/abs/2408.13459