STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910773729558528 |
|---|---|
| author | Xie, Rui Liu, Yinhong Zhou, Penghao Zhao, Chen Zhou, Jun Zhang, Kai Zhang, Zhenyu Yang, Jian Yang, Zhenheng Tai, Ying |
| author_facet | Xie, Rui Liu, Yinhong Zhou, Penghao Zhao, Chen Zhou, Jun Zhang, Kai Zhang, Zhenyu Yang, Jian Yang, Zhenheng Tai, Ying |
| contents | Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_02976 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution Xie, Rui Liu, Yinhong Zhou, Penghao Zhao, Chen Zhou, Jun Zhang, Kai Zhang, Zhenyu Yang, Jian Yang, Zhenheng Tai, Ying Computer Vision and Pattern Recognition Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets. |
| title | STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2501.02976 |