RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Weisong, Zhou, Jingkai, Zhu, Xiangyu, Chen, Weihua, Zhang, Xiao-Yu, Lei, Zhen, Wang, Fan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918104289771520
author Zhao, Weisong
Zhou, Jingkai
Zhu, Xiangyu
Chen, Weihua
Zhang, Xiao-Yu
Lei, Zhen
Wang, Fan
author_facet Zhao, Weisong
Zhou, Jingkai
Zhu, Xiangyu
Chen, Weihua
Zhang, Xiao-Yu
Lei, Zhen
Wang, Fan
contents Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite recent advances, three critical challenges persist in VSR community: 1) Inconsistent modeling of temporal dynamics in foundational models; 2) limited high-frequency detail recovery under complex real-world degradations; and 3) insufficient evaluation of detail enhancement and 4K super-resolution, as current methods primarily rely on 720P datasets with inadequate details. To address these challenges, we propose RealisVSR, a high-frequency detail-enhanced video diffusion model with three core innovations: 1) Consistency Preserved ControlNet (CPC) architecture integrated with the Wan2.1 video diffusion to model the smooth and complex motions and suppress artifacts; 2) High-Frequency Rectified Diffusion Loss (HR-Loss) combining wavelet decomposition and HOG feature constraints for texture restoration; 3) RealisVideo-4K, the first public 4K VSR benchmark containing 1,000 high-definition video-text pairs. Leveraging the advanced spatio-temporal guidance of Wan2.1, our method requires only 5-25% of the training data volume compared to existing approaches. Extensive experiments on VSR benchmarks (REDS, SPMCS, UDM10, YouTube-HQ, VideoLQ, RealisVideo-720P) demonstrate our superiority, particularly in ultra-high-resolution scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution
Zhao, Weisong
Zhou, Jingkai
Zhu, Xiangyu
Chen, Weihua
Zhang, Xiao-Yu
Lei, Zhen
Wang, Fan
Image and Video Processing
Computer Vision and Pattern Recognition
Video Super-Resolution (VSR) has achieved significant progress through diffusion models, effectively addressing the over-smoothing issues inherent in GAN-based methods. Despite recent advances, three critical challenges persist in VSR community: 1) Inconsistent modeling of temporal dynamics in foundational models; 2) limited high-frequency detail recovery under complex real-world degradations; and 3) insufficient evaluation of detail enhancement and 4K super-resolution, as current methods primarily rely on 720P datasets with inadequate details. To address these challenges, we propose RealisVSR, a high-frequency detail-enhanced video diffusion model with three core innovations: 1) Consistency Preserved ControlNet (CPC) architecture integrated with the Wan2.1 video diffusion to model the smooth and complex motions and suppress artifacts; 2) High-Frequency Rectified Diffusion Loss (HR-Loss) combining wavelet decomposition and HOG feature constraints for texture restoration; 3) RealisVideo-4K, the first public 4K VSR benchmark containing 1,000 high-definition video-text pairs. Leveraging the advanced spatio-temporal guidance of Wan2.1, our method requires only 5-25% of the training data volume compared to existing approaches. Extensive experiments on VSR benchmarks (REDS, SPMCS, UDM10, YouTube-HQ, VideoLQ, RealisVideo-720P) demonstrate our superiority, particularly in ultra-high-resolution scenarios.
title RealisVSR: Detail-enhanced Diffusion for Real-World 4K Video Super-Resolution
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19138