Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhikai, Long, Fuchen, Qiu, Zhaofan, Yao, Ting, Zhou, Wengang, Luo, Jiebo, Mei, Tao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913282510553088
author Chen, Zhikai
Long, Fuchen
Qiu, Zhaofan
Yao, Ting
Zhou, Wengang
Luo, Jiebo
Mei, Tao
author_facet Chen, Zhikai
Long, Fuchen
Qiu, Zhaofan
Yao, Ting
Zhou, Wengang
Luo, Jiebo
Mei, Tao
contents Diffusion models are just at a tipping point for image super-resolution task. Nevertheless, it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of visual appearance from low-resolution to high-resolution videos, but also the temporal consistency across video frames. In this paper, we propose a novel approach, pursuing Spatial Adaptation and Temporal Coherence (SATeCo), for video super-resolution. SATeCo pivots on learning spatial-temporal guidance from low-resolution videos to calibrate both latent-space high-resolution video denoising and pixel-space video reconstruction. Technically, SATeCo freezes all the parameters of the pre-trained UNet and VAE, and only optimizes two deliberately-designed spatial feature adaptation (SFA) and temporal feature alignment (TFA) modules, in the decoder of UNet and VAE. SFA modulates frame features via adaptively estimating affine parameters for each pixel, guaranteeing pixel-wise guidance for high-resolution frame synthesis. TFA delves into feature interaction within a 3D local window (tubelet) through self-attention, and executes cross-attention between tubelet and its low-resolution counterpart to guide temporal feature alignment. Extensive experiments conducted on the REDS4 and Vid4 datasets demonstrate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17000
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
Chen, Zhikai
Long, Fuchen
Qiu, Zhaofan
Yao, Ting
Zhou, Wengang
Luo, Jiebo
Mei, Tao
Computer Vision and Pattern Recognition
Multimedia
Diffusion models are just at a tipping point for image super-resolution task. Nevertheless, it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of visual appearance from low-resolution to high-resolution videos, but also the temporal consistency across video frames. In this paper, we propose a novel approach, pursuing Spatial Adaptation and Temporal Coherence (SATeCo), for video super-resolution. SATeCo pivots on learning spatial-temporal guidance from low-resolution videos to calibrate both latent-space high-resolution video denoising and pixel-space video reconstruction. Technically, SATeCo freezes all the parameters of the pre-trained UNet and VAE, and only optimizes two deliberately-designed spatial feature adaptation (SFA) and temporal feature alignment (TFA) modules, in the decoder of UNet and VAE. SFA modulates frame features via adaptively estimating affine parameters for each pixel, guaranteeing pixel-wise guidance for high-resolution frame synthesis. TFA delves into feature interaction within a 3D local window (tubelet) through self-attention, and executes cross-attention between tubelet and its low-resolution counterpart to guide temporal feature alignment. Extensive experiments conducted on the REDS4 and Vid4 datasets demonstrate the effectiveness of our approach.
title Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2403.17000