OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhiqiang, Sun, Zhaomang, Zhou, Tong, Fu, Bingtao, Cong, Ji, Dong, Yitong, Zhang, Huaqi, Tang, Xuan, Chen, Mingsong, Wei, Xian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914167926030336
author Wu, Zhiqiang
Sun, Zhaomang
Zhou, Tong
Fu, Bingtao
Cong, Ji
Dong, Yitong
Zhang, Huaqi
Tang, Xuan
Chen, Mingsong
Wei, Xian
author_facet Wu, Zhiqiang
Sun, Zhaomang
Zhou, Tong
Fu, Bingtao
Cong, Ji
Dong, Yitong
Zhang, Huaqi
Tang, Xuan
Chen, Mingsong
Wei, Xian
contents Denoising Diffusion Probabilistic Models (DDPMs) show promising potential in one-step Real-World Image Super-Resolution (Real-ISR). Current one-step Real-ISR methods typically inject the low-quality (LQ) image latent representation at the start or end timestep of the DDPM scheduler. Recent studies have begun to note that the LQ image latent and the pre-trained noisy latent representations are intuitively closer at a mid-timestep. However, a quantitative analysis of these latent representations remains lacking. Considering these latent representations can be decomposed into signal and noise, we propose a method based on the Signal-to-Noise Ratio (SNR) to pre-compute an average optimal mid-timestep for injection. To better approximate the pre-trained noisy latent representation, we further introduce the Latent Representation Refinement (LRR) loss via a LoRA-enhanced VAE encoder. We also fine-tune the backbone of the DDPM-based generative model using LoRA to perform one-step denoising at the average optimal mid-timestep. Based on these components, we present OMGSR, a GAN-based Real-ISR framework that employs a DDPM-based generative model as the generator and a DINOv3-ConvNeXt model with multi-level discriminator heads as the discriminator. We also propose the DINOv3-ConvNeXt DISTS (Dv3CD) loss, which is enhanced for structural perception at varying resolutions. Within the OMGSR framework, we develop OMGSR-S based on SD2.1-base. An ablation study confirms that our pre-computation strategy and LRR loss significantly improve the baseline. Comparative studies demonstrate that OMGSR-S achieves state-of-the-art performance across multiple metrics. Code is available at \hyperlink{Github}{https://github.com/wuer5/OMGSR}.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
Wu, Zhiqiang
Sun, Zhaomang
Zhou, Tong
Fu, Bingtao
Cong, Ji
Dong, Yitong
Zhang, Huaqi
Tang, Xuan
Chen, Mingsong
Wei, Xian
Computer Vision and Pattern Recognition
Artificial Intelligence
Denoising Diffusion Probabilistic Models (DDPMs) show promising potential in one-step Real-World Image Super-Resolution (Real-ISR). Current one-step Real-ISR methods typically inject the low-quality (LQ) image latent representation at the start or end timestep of the DDPM scheduler. Recent studies have begun to note that the LQ image latent and the pre-trained noisy latent representations are intuitively closer at a mid-timestep. However, a quantitative analysis of these latent representations remains lacking. Considering these latent representations can be decomposed into signal and noise, we propose a method based on the Signal-to-Noise Ratio (SNR) to pre-compute an average optimal mid-timestep for injection. To better approximate the pre-trained noisy latent representation, we further introduce the Latent Representation Refinement (LRR) loss via a LoRA-enhanced VAE encoder. We also fine-tune the backbone of the DDPM-based generative model using LoRA to perform one-step denoising at the average optimal mid-timestep. Based on these components, we present OMGSR, a GAN-based Real-ISR framework that employs a DDPM-based generative model as the generator and a DINOv3-ConvNeXt model with multi-level discriminator heads as the discriminator. We also propose the DINOv3-ConvNeXt DISTS (Dv3CD) loss, which is enhanced for structural perception at varying resolutions. Within the OMGSR framework, we develop OMGSR-S based on SD2.1-base. An ablation study confirms that our pre-computation strategy and LRR loss significantly improve the baseline. Comparative studies demonstrate that OMGSR-S achieves state-of-the-art performance across multiple metrics. Code is available at \hyperlink{Github}{https://github.com/wuer5/OMGSR}.
title OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.08227