EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Haizhen, Du, Kunpeng, Yan, Qiangyu, Lu, Sen, Han, Jianhong, Chen, Hanting, Hu, Hailin, Hu, Jie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913108533968896
author Xie, Haizhen
Du, Kunpeng
Yan, Qiangyu
Lu, Sen
Han, Jianhong
Chen, Hanting
Hu, Hailin
Hu, Jie
author_facet Xie, Haizhen
Du, Kunpeng
Yan, Qiangyu
Lu, Sen
Han, Jianhong
Chen, Hanting
Hu, Hailin
Hu, Jie
contents Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally relied on U-Net architectures, recent advancements have demonstrated that Diffusion Transformers (DiT) achieve significantly higher performance in this domain. In this work, we introduce Enhancing Anything Model (EAM), a novel BSR method that leverages DiT and outperforms previous U-Net-based approaches. We introduce a novel block, $Ψ$-DiT, which effectively guides the DiT to enhance image restoration. This block employs a low-resolution latent as a separable flow injection control, forming a triple-flow architecture that effectively leverages the prior knowledge embedded in the pre-trained DiT. To fully exploit the prior guidance capabilities of T2I models and enhance their generalization in BSR, we introduce a progressive Masked Image Modeling strategy, which also reduces training costs. Additionally, we propose a subject-aware prompt generation strategy that employs a robust multi-modal model in an in-context learning framework. This strategy automatically identifies key image areas, provides detailed descriptions, and optimizes the utilization of T2I diffusion priors. Our experiments demonstrate that EAM achieves state-of-the-art results across multiple datasets, outperforming existing methods in both quantitative metrics and visual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05209
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
Xie, Haizhen
Du, Kunpeng
Yan, Qiangyu
Lu, Sen
Han, Jianhong
Chen, Hanting
Hu, Hailin
Hu, Jie
Computer Vision and Pattern Recognition
Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally relied on U-Net architectures, recent advancements have demonstrated that Diffusion Transformers (DiT) achieve significantly higher performance in this domain. In this work, we introduce Enhancing Anything Model (EAM), a novel BSR method that leverages DiT and outperforms previous U-Net-based approaches. We introduce a novel block, $Ψ$-DiT, which effectively guides the DiT to enhance image restoration. This block employs a low-resolution latent as a separable flow injection control, forming a triple-flow architecture that effectively leverages the prior knowledge embedded in the pre-trained DiT. To fully exploit the prior guidance capabilities of T2I models and enhance their generalization in BSR, we introduce a progressive Masked Image Modeling strategy, which also reduces training costs. Additionally, we propose a subject-aware prompt generation strategy that employs a robust multi-modal model in an in-context learning framework. This strategy automatically identifies key image areas, provides detailed descriptions, and optimizes the utilization of T2I diffusion priors. Our experiments demonstrate that EAM achieves state-of-the-art results across multiple datasets, outperforming existing methods in both quantitative metrics and visual quality.
title EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.05209