One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Selikhanovych, Daniil, Li, David, Leonov, Aleksei, Gushchin, Nikita, Kushneriuk, Sergei, Filippov, Alexander, Burnaev, Evgeny, Koshelev, Iaroslav, Korotin, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911420100116480
author Selikhanovych, Daniil
Li, David
Leonov, Aleksei
Gushchin, Nikita
Kushneriuk, Sergei
Filippov, Alexander
Burnaev, Evgeny
Koshelev, Iaroslav
Korotin, Alexander
author_facet Selikhanovych, Daniil
Li, David
Leonov, Aleksei
Gushchin, Nikita
Kushneriuk, Sergei
Filippov, Alexander
Burnaev, Evgeny
Koshelev, Iaroslav
Korotin, Alexander
contents Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of several methods to accelerate diffusion-based SR models, some (e.g., SinSR) fail to produce realistic perceptual details, while others (e.g., OSEDiff) may hallucinate non-existent structures. To overcome these issues, we present RSD, a new distillation method for ResShift. Our method is based on training the student network to produce images such that a new fake ResShift model trained on them will coincide with the teacher model. RSD achieves single-step restoration and outperforms the teacher by a noticeable margin in various perceptual metrics (LPIPS, CLIPIQA, MUSIQ). We show that our distillation method can surpass SinSR, the other distillation-based method for ResShift, making it on par with state-of-the-art diffusion SR distillation methods with limited computational costs in terms of perceptual quality. Compared to SR methods based on pre-trained text-to-image models, RSD produces competitive perceptual quality and requires fewer parameters, GPU memory, and training cost. We provide experimental results on various real-world and synthetic datasets, including RealSR, RealSet65, DRealSR, ImageNet, and DIV2K.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation
Selikhanovych, Daniil
Li, David
Leonov, Aleksei
Gushchin, Nikita
Kushneriuk, Sergei
Filippov, Alexander
Burnaev, Evgeny
Koshelev, Iaroslav
Korotin, Alexander
Computer Vision and Pattern Recognition
Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of several methods to accelerate diffusion-based SR models, some (e.g., SinSR) fail to produce realistic perceptual details, while others (e.g., OSEDiff) may hallucinate non-existent structures. To overcome these issues, we present RSD, a new distillation method for ResShift. Our method is based on training the student network to produce images such that a new fake ResShift model trained on them will coincide with the teacher model. RSD achieves single-step restoration and outperforms the teacher by a noticeable margin in various perceptual metrics (LPIPS, CLIPIQA, MUSIQ). We show that our distillation method can surpass SinSR, the other distillation-based method for ResShift, making it on par with state-of-the-art diffusion SR distillation methods with limited computational costs in terms of perceptual quality. Compared to SR methods based on pre-trained text-to-image models, RSD produces competitive perceptual quality and requires fewer parameters, GPU memory, and training cost. We provide experimental results on various real-world and synthetic datasets, including RealSR, RealSet65, DRealSR, ImageNet, and DIV2K.
title One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.13358