STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Rui, Liu, Yinhong, Zhou, Penghao, Zhao, Chen, Zhou, Jun, Zhang, Kai, Zhang, Zhenyu, Yang, Jian, Yang, Zhenheng, Tai, Ying
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910773729558528
author Xie, Rui
Liu, Yinhong
Zhou, Penghao
Zhao, Chen
Zhou, Jun
Zhang, Kai
Zhang, Zhenyu
Yang, Jian
Yang, Zhenheng
Tai, Ying
author_facet Xie, Rui
Liu, Yinhong
Zhou, Penghao
Zhao, Chen
Zhou, Jun
Zhang, Kai
Zhang, Zhenyu
Yang, Jian
Yang, Zhenheng
Tai, Ying
contents Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02976
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
Xie, Rui
Liu, Yinhong
Zhou, Penghao
Zhao, Chen
Zhou, Jun
Zhang, Kai
Zhang, Zhenyu
Yang, Jian
Yang, Zhenheng
Tai, Ying
Computer Vision and Pattern Recognition
Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets.
title STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.02976