MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xinrui, Zhang, Jinrong, Wu, Jianlong, Chen, Chong, Nie, Liqiang, Lin, Zhouchen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908685330022400
author Li, Xinrui
Zhang, Jinrong
Wu, Jianlong
Chen, Chong
Nie, Liqiang
Lin, Zhouchen
author_facet Li, Xinrui
Zhang, Jinrong
Wu, Jianlong
Chen, Chong
Nie, Liqiang
Lin, Zhouchen
contents Text-to-image (T2I) models have ushered in a new era of real-world image super-resolution (Real-ISR) due to their rich internal implicit knowledge for multimodal learning. Although bringing high-level semantic priors and dense pixel guidance have led to advances in reconstruction, we identified several critical phenomena by analyzing the behavior of existing T2I-based Real-ISR methods: (1) Fine detail deficiency, which ultimately leads to incorrect reconstruction in local regions. (2) Block-wise semantic inconsistency, which results in distracted semantic interpretations across U-Net blocks. (3) Edge ambiguity, which causes noticeable structural degradation. Building upon these observations, we first introduce MegaSR, which enhances the T2I-based Real-ISR models with fine-grained customized semantics and expressive guidance to unlock semantically rich and structurally consistent reconstruction. Then, we propose the Customized Semantics Module (CSM) to supplement fine-grained semantics from the image modality and regulate the semantic fusion between multi-level knowledge to realize customization for different U-Net blocks. Besides the semantic adaptation, we identify expressive multimodal signals through pair-wise comparisons and introduce the Multimodal Signal Fusion Module (MSFM) to aggregate them for structurally consistent reconstruction. Extensive experiments on real-world and synthetic datasets demonstrate the superiority of the method. Notably, it not only achieves state-of-the-art performance on quality-driven metrics but also remains competitive on fidelity-focused metrics, striking a balance between perceptual realism and faithful content reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution
Li, Xinrui
Zhang, Jinrong
Wu, Jianlong
Chen, Chong
Nie, Liqiang
Lin, Zhouchen
Computer Vision and Pattern Recognition
Text-to-image (T2I) models have ushered in a new era of real-world image super-resolution (Real-ISR) due to their rich internal implicit knowledge for multimodal learning. Although bringing high-level semantic priors and dense pixel guidance have led to advances in reconstruction, we identified several critical phenomena by analyzing the behavior of existing T2I-based Real-ISR methods: (1) Fine detail deficiency, which ultimately leads to incorrect reconstruction in local regions. (2) Block-wise semantic inconsistency, which results in distracted semantic interpretations across U-Net blocks. (3) Edge ambiguity, which causes noticeable structural degradation. Building upon these observations, we first introduce MegaSR, which enhances the T2I-based Real-ISR models with fine-grained customized semantics and expressive guidance to unlock semantically rich and structurally consistent reconstruction. Then, we propose the Customized Semantics Module (CSM) to supplement fine-grained semantics from the image modality and regulate the semantic fusion between multi-level knowledge to realize customization for different U-Net blocks. Besides the semantic adaptation, we identify expressive multimodal signals through pair-wise comparisons and introduce the Multimodal Signal Fusion Module (MSFM) to aggregate them for structurally consistent reconstruction. Extensive experiments on real-world and synthetic datasets demonstrate the superiority of the method. Notably, it not only achieves state-of-the-art performance on quality-driven metrics but also remains competitive on fidelity-focused metrics, striking a balance between perceptual realism and faithful content reconstruction.
title MegaSR: Mining Customized Semantics and Expressive Guidance for Real-World Image Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.08096