XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qu, Yunpeng, Yuan, Kun, Zhao, Kai, Xie, Qizhi, Hao, Jinhua, Sun, Ming, Zhou, Chao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913437668343808
author Qu, Yunpeng
Yuan, Kun
Zhao, Kai
Xie, Qizhi
Hao, Jinhua
Sun, Ming
Zhou, Chao
author_facet Qu, Yunpeng
Yuan, Kun
Zhao, Kai
Xie, Qizhi
Hao, Jinhua
Sun, Ming
Zhou, Chao
contents Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation information, resulting in restoration images with incorrect content or unrealistic artifacts. To address these issues, we propose a \textit{Cross-modal Priors for Super-Resolution (XPSR)} framework. Within XPSR, to acquire precise and comprehensive semantic conditions for the diffusion model, cutting-edge Multimodal Large Language Models (MLLMs) are utilized. To facilitate better fusion of cross-modal priors, a \textit{Semantic-Fusion Attention} is raised. To distill semantic-preserved information instead of undesired degradations, a \textit{Degradation-Free Constraint} is attached between LR and its high-resolution (HR) counterpart. Quantitative and qualitative results show that XPSR is capable of generating high-fidelity and high-realism images across synthetic and real-world datasets. Codes are released at \url{https://github.com/qyp2000/XPSR}.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution
Qu, Yunpeng
Yuan, Kun
Zhao, Kai
Xie, Qizhi
Hao, Jinhua
Sun, Ming
Zhou, Chao
Computer Vision and Pattern Recognition
Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation information, resulting in restoration images with incorrect content or unrealistic artifacts. To address these issues, we propose a \textit{Cross-modal Priors for Super-Resolution (XPSR)} framework. Within XPSR, to acquire precise and comprehensive semantic conditions for the diffusion model, cutting-edge Multimodal Large Language Models (MLLMs) are utilized. To facilitate better fusion of cross-modal priors, a \textit{Semantic-Fusion Attention} is raised. To distill semantic-preserved information instead of undesired degradations, a \textit{Degradation-Free Constraint} is attached between LR and its high-resolution (HR) counterpart. Quantitative and qualitative results show that XPSR is capable of generating high-fidelity and high-realism images across synthetic and real-world datasets. Codes are released at \url{https://github.com/qyp2000/XPSR}.
title XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.05049