One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911420100116480 |
|---|---|
| author | Selikhanovych, Daniil Li, David Leonov, Aleksei Gushchin, Nikita Kushneriuk, Sergei Filippov, Alexander Burnaev, Evgeny Koshelev, Iaroslav Korotin, Alexander |
| author_facet | Selikhanovych, Daniil Li, David Leonov, Aleksei Gushchin, Nikita Kushneriuk, Sergei Filippov, Alexander Burnaev, Evgeny Koshelev, Iaroslav Korotin, Alexander |
| contents | Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of several methods to accelerate diffusion-based SR models, some (e.g., SinSR) fail to produce realistic perceptual details, while others (e.g., OSEDiff) may hallucinate non-existent structures. To overcome these issues, we present RSD, a new distillation method for ResShift. Our method is based on training the student network to produce images such that a new fake ResShift model trained on them will coincide with the teacher model. RSD achieves single-step restoration and outperforms the teacher by a noticeable margin in various perceptual metrics (LPIPS, CLIPIQA, MUSIQ). We show that our distillation method can surpass SinSR, the other distillation-based method for ResShift, making it on par with state-of-the-art diffusion SR distillation methods with limited computational costs in terms of perceptual quality. Compared to SR methods based on pre-trained text-to-image models, RSD produces competitive perceptual quality and requires fewer parameters, GPU memory, and training cost. We provide experimental results on various real-world and synthetic datasets, including RealSR, RealSet65, DRealSR, ImageNet, and DIV2K. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_13358 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation Selikhanovych, Daniil Li, David Leonov, Aleksei Gushchin, Nikita Kushneriuk, Sergei Filippov, Alexander Burnaev, Evgeny Koshelev, Iaroslav Korotin, Alexander Computer Vision and Pattern Recognition Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of several methods to accelerate diffusion-based SR models, some (e.g., SinSR) fail to produce realistic perceptual details, while others (e.g., OSEDiff) may hallucinate non-existent structures. To overcome these issues, we present RSD, a new distillation method for ResShift. Our method is based on training the student network to produce images such that a new fake ResShift model trained on them will coincide with the teacher model. RSD achieves single-step restoration and outperforms the teacher by a noticeable margin in various perceptual metrics (LPIPS, CLIPIQA, MUSIQ). We show that our distillation method can surpass SinSR, the other distillation-based method for ResShift, making it on par with state-of-the-art diffusion SR distillation methods with limited computational costs in terms of perceptual quality. Compared to SR methods based on pre-trained text-to-image models, RSD produces competitive perceptual quality and requires fewer parameters, GPU memory, and training cost. We provide experimental results on various real-world and synthetic datasets, including RealSR, RealSet65, DRealSR, ImageNet, and DIV2K. |
| title | One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.13358 |