Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Yiwen, Liang, Ying, Zhang, Yuxuan, Chai, Xinning, Cheng, Zhengxue, Qin, Yingsheng, Yang, Yucai, Xie, Rong, Song, Li
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908317608050688
author Wang, Yiwen
Liang, Ying
Zhang, Yuxuan
Chai, Xinning
Cheng, Zhengxue
Qin, Yingsheng
Yang, Yucai
Xie, Rong
Song, Li
author_facet Wang, Yiwen
Liang, Ying
Zhang, Yuxuan
Chai, Xinning
Cheng, Zhengxue
Qin, Yingsheng
Yang, Yucai
Xie, Rong
Song, Li
contents Due to the disparity between real-world degradations in user-generated content(UGC) images and synthetic degradations, traditional super-resolution methods struggle to generalize effectively, necessitating a more robust approach to model real-world distortions. In this paper, we propose a novel approach to UGC image super-resolution by integrating semantic guidance into a diffusion framework. Our method addresses the inconsistency between degradations in wild and synthetic datasets by separately simulating the degradation processes on the LSDIR dataset and combining them with the official paired training set. Furthermore, we enhance degradation removal and detail generation by incorporating a pretrained semantic extraction model (SAM2) and fine-tuning key hyperparameters for improved perceptual fidelity. Extensive experiments demonstrate the superiority of our approach against state-of-the-art methods. Additionally, the proposed model won second place in the CVPR NTIRE 2025 Short-form UGC Image Super-Resolution Challenge, further validating its effectiveness. The code is available at https://github.c10pom/Moonsofang/NTIRE-2025-SRlab.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution
Wang, Yiwen
Liang, Ying
Zhang, Yuxuan
Chai, Xinning
Cheng, Zhengxue
Qin, Yingsheng
Yang, Yucai
Xie, Rong
Song, Li
Computer Vision and Pattern Recognition
Due to the disparity between real-world degradations in user-generated content(UGC) images and synthetic degradations, traditional super-resolution methods struggle to generalize effectively, necessitating a more robust approach to model real-world distortions. In this paper, we propose a novel approach to UGC image super-resolution by integrating semantic guidance into a diffusion framework. Our method addresses the inconsistency between degradations in wild and synthetic datasets by separately simulating the degradation processes on the LSDIR dataset and combining them with the official paired training set. Furthermore, we enhance degradation removal and detail generation by incorporating a pretrained semantic extraction model (SAM2) and fine-tuning key hyperparameters for improved perceptual fidelity. Extensive experiments demonstrate the superiority of our approach against state-of-the-art methods. Additionally, the proposed model won second place in the CVPR NTIRE 2025 Short-form UGC Image Super-Resolution Challenge, further validating its effectiveness. The code is available at https://github.c10pom/Moonsofang/NTIRE-2025-SRlab.
title Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.09887